Slashdot Mirror


Cooling Challenges an Issue In Rackspace Outage

miller60 writes "If your data center's cooling system fails, how long do you have before your servers overheat? The shrinking window for recovery from a grid power outage appears to have been an issue in Monday night's downtime for some customers of Rackspace, which has historically been among the most reliable hosting providers. The company's Dallas data center lost power when a traffic accident damaged a nearby power transformer. There were difficulties getting the chillers fully back online (it's not clear if this was equipment issues or subsequent power bumps) and temperatures rose in the data center, forcing Rackspace to take customer servers offline to protect the equipment. A recent study found that a data center running at 5 kilowatts per server cabinet may experience a thermal shutdown in as little as three minutes during a power outage. The short recovery window from cooling outages has been a hot topic in discussions of data center energy efficiency. One strategy being actively debated is raising the temperature set point in the data center, which trims power bills but may create a less forgiving environment in a cooling outage."

3 of 294 comments (clear)

  1. This is number 3 by DuctTape · · Score: 5, Informative
    This is actually Rackspace's number 3 outage in the past couple days. My company was only (!) affected by outages 1 and 2. My boss would have had a fit if number 3 would have taken us down for the third time.

    Other publications have noted it was number 3, too.

    DT

    --
    Is this thing on? Hello?
  2. Re:How to estimate the cooling needs? by CaptainPatent · · Score: 4, Informative

    Is there a general rule for figuring out how many BTUs of cooling you need for a given wattage of power supplies? I actually found a good article about this earlier on and it helped me purchase a window unit for a closet turned server-room. Hope that helps out a bit.
    --
    Well, back to rejecting software patent applications.
  3. Re:Which only shows by jandrese · · Score: 4, Informative

    If you want 100% uptime (which is impossible, but you can put enough 9s in your reliability to be close enough), you need to have your data distributed across multiple data centers, geographically separate, and over provisioned enough that the loss of one data center won't cause the others to be overloaded. It's important to keep your geographical separation large because you never know when the entire eastern (or western) seaboard will experience complete power failure or when a major backhaul router will go down/have a line cut. Preferably each data center should get power from multiple sources if they can, and multiple POPs on the internet from each center is almost mandatory.

    --

    I read the internet for the articles.