Showing posts with label cooling. Show all posts
Showing posts with label cooling. Show all posts

Thursday, July 16, 2009

Adding a Geographic Element to PUE Calculations

The PUE metric has become one of the most significant metrics for measuring the gross efficiency of a data center. As data center operators boast of PUE numbers that approach the optimal rating of 1.0, it's often difficult to separate out environmental or regional factors.

Is a PUE of 1.5 in Phoenix better or worse than a PUE of 1.4 in Seattle?

It depends. In absolute numbers, the lower PUE provides an indicator of the most efficient facility. However, achieving a PUE of 1.5 in Phoenix is much more difficult than an equivalent or slightly lower number in Seattle because Phoenix is so much hotter and requires more air conditioning. Moving data centers to cooler locations helps the PUE rating, but sometimes data centers need to be located in a specific city or region. How can you compare PUE values in regions with different environmental conditions?

One possible approach is to add a geographic compensating factor:

gPUE = G * PUE

The geographic compensating factor G would be determined by The Green Grid or other trusted body based on compiled weather data. Ideally, this could be calculated empirically through a formula using data maintained by the U.S. Department of Energy (refer to this blog link for information on that data and a free tool to visually represent that data).

This approach would allow somebody to measure the technical innovation of a given facility while providing an adjustment to account for geographic disparities in temperature, wind, solar loading, etc. It's not a perfect solution (since some cooling optimizations might not work in cooler or hotter climates), but it provides some measure of equalization to facilitate more equitable comparisons between PUE claims in different locations.

--kb

Monday, June 15, 2009

Making Ice to Lower PUE and TCO

At night, demand on the grid is lower, energy costs tend to be lower, and temperatures are also lower. These three factors make night an attractive time to produce thermal storage. This allows facility managers to time-shift HVAC-related energy costs to reduce peak demands on the grid and lower energy costs.

Although some facility managers have developed their own methods for time-shifting HVAC energy requirements, Ice Energy may be the first vendor to market a product specifically designed to do this. The Ice Bear* distributed energy storage system provides up to 5 tons of cooling load during peak hours.

It's good to see innovative products like this coming to market.

--kb

Monday, May 11, 2009

How a Good Metric Could Drive Bad Behaviors

The PUE (Power Usage Effectiveness) metric from The Green Grid has become a widely referenced benchmark in the data center community, and justifiably so. However, there can be a dark side to following this metric blindly.


Introduction

PUE is defined as follows:

PUE = Total Facility Power/IT Equipment Power

Using the PUE metric, a facility manager can judge what ratio of power is lost in "overhead" (infrastructure) to operate the facility. A PUE of 1.6 to 2.0 is typical, but facility managers are striving to approach a PUE of 1.0, the idealized state.

Companies willing to drive more sustainable practices may incent facility managers to improve facility PUE levels. However, if this is done without context towards the overall energy or other resource consumption, it could drive inefficient behaviors.



Issue #1: Dissimilar Infrastructure Power Scaling

If a facility manager tracks PUE over a variety of workloads, they will see how the data center's infrastructure power consumption tracks with the IT load. Ideally, the infrastructure overhead (HVAC system, UPS system, etc.) will match linearly with the consumption of the servers and other gear in the data center, but this is rarely the case.



In many cases, the fixed overhead for power and cooling systems will become a higher percentage of overall power consumption as the IT load diminished. In other cases, there will be significant step functions in overall power consumption as large infrastructure items such as chillers, CRACs, or other equipment is turned on or off (as depicted in the graph to the left).

In such situations, reducing the IT power consumption could increase the PUE even if it reduces the overall energy consumption of the data center. People will often act in the direction towards which they are incented (i.e., what improves their paycheck). Managers incented to improve PUE without any clear tie-in to overall energy consumption might be reluctant to shut off unused servers or aggressively implement power saving features on their IT infrastructure if it increased their PUE--even if doing so would reduce overall facility power consumption.

Ensuring overall energy consumption is part of the incentive package (not just PUE) is critical to driving the desired behaviors.

[Part of this needs to be linked with overall productivity of the data center so that increased use of the data centers is encouraged while still incenting improved efficiency. I'll write about this in an upcoming post.]



Issue #2: Shifting Infrastructure Loads to IT

Another issue to watch is a desire to classify some infrastructure-like services as IT loads in order to improve PUE efficiencies. Examples of this include moving UPS systems into IT racks or putting large air-mover devices into equipment cabinets and trying to classify them as IT loads. This is "gaming" the system and should be actively discouraged.

The Green Grid is aware of this issue and is adding more guidelines to help people improve the accuracy and consistency of their PUE reporting.



Issue #3: Improving Infrastructure Efficiency at the Expense of IT

The third issue to watch is a move towards facility or equipment practices that reduce the infrastructure power consumption but increase the IT power consumption. In particular, the adoption of higher operating temperatures for data centers warrants particular scrutiny.

I've noted previously that there are significant gains possible by raising data center temperatures and making greater use of dry-side or wet-side economizers. However, it's important to compare the energy savings on the infrastructure side with the energy costs on the IT side. At higher temperatures, leakage currents in silicon increase and fans inside servers need to run faster to move more air through each server.

Increase the IT consumption and lower the infrastructure consumption and you get a two-fer: the PUE numerator goes down and the PUE denominator goes up, lowering the overall PUE. However, if the net power consumption doesn't go down, it usually** doesn't make sense to increase the ambient temperature. Once again, looking at overall power consumption in addition to PUE is important in incenting the proper behaviors.

--kb


**Note: For greenfield (new) data centers or substantial datacenter retrofits, raising the allowed data center temperature may eliminate or substantially reduce the CapEx (capital expenditure) cost for that data center even if the direct energy costs are slightly higher. For example, if a data center doesn't need to purchase a chiller unit, that could shave millions of dollars off the construction cost for a facility. In such cases, more complicated parameters will be needed to evaluate the benefits of raising the ambient temperature in the facility; these likely will include a net present value analysis for the CapEx savings vs. OpEx (operating expense) costs, consideration of real estate savings, etc. The real win is when both CapEx costs are avoided AND OpEx costs are lower.

Wednesday, April 29, 2009

Human Side of Higher Data Center Temperatures

With all the talk of hotter data center temperatures, one item that has often been overlooked is what happens to the poor soul tasked with going in and servicing equipment in that data center. Imagine having to work in a facility at 40°C (104°F) for several hours at a time--and that's at the equipment input. The exhaust temperature on the back side of the rack could easily be 55°C (131°F).

One approach is to adopt a "fail in place" model where technicians never go into a production facility, but even Google has technicians adding and replacing individual servers in their containerized data centers.

Other approaches to consider:
  • Localized spot cooling. A very small air conditioner could take the edge off the area in front of a rack.
  • Perform service operations at night or when it's reasonably cool.

This last suggestion may seem too simplistic at first, but it's actually quite practical. In a facility with sufficient redundancy to ensure high availability, server replacement should be able to wait up to 24 hours. Operating a data center at consistently high temperatures will end up increasing power consumption in the IT equipment. It only makes sense to use higher temperatures in a data center when using optimizers to eliminate or substantially reduce HVAC CapEx and OpEx costs.

If a data center is using economizers, the temperature in the data center should drop when the outside temperature drops. Even in relatively warm areas during summer months, there are substantial times each day where the temperature drops to reasonable levels in which technicians can comfortably work.

--kb

Monday, April 13, 2009

NEBS vs. the Hottest Place on Earth

As mentioned in Higher Temperatures for Data Center and Processors for Higher Temps, various groups are pushing for higher and higher ambient temperatures in data centers. At Google's Efficient Data Center Summit last week, Amazon's James Hamilton brought up an interesting point in his slides and blog about ambient temperatures:
the hottest place on earth over recorded history was Al Aziziyah Libya in 1922 where 136F (58C) was indicated

James went on to note during his talk that telecommunications equipment designed to the NEBS (Network Equipment Building System) standards routinely has to handle temperatures up to 40°C.

Actually, the story is better than that. NEBS-GR-63 (the key NEBS specification dealing with environmental conditions for equipment in telecommunications central offices) requires equipment to handle 40°C long-term ambient temperatures, but telecommunications equipment certified at the shelf (chassis) level needs to be able to operate at 55°C ambient for up to 96 hours at a time and up to 360 hours per year [the 360 hours is for reliability calculations]. This means that much of the NEBS-rated equipment for data centers can operate at temperatures that are only 3°C lower than the highest natural temperature ever recorded on Earth, as noted by James.

Given the common engineering penchant to provide some guardband on products vs. the official specifications, even a 58°C ambient is not out of the question. This means that NEBS-rated equipment could be good candidates for data centers operating at high temperatures.

But can you get decent performance in NEBS-rated servers? Yes! For example, vendors such as Radisys, Kontron, and Emerson have announced blade servers with Intel's new 5500 (aka "Nehalem") processors, and their bladed servers commonly are NEBS certified to operate at 55°C. This would allow the latest server technology to operate in the most demanding environments.

--kb

Friday, February 13, 2009

Processors for Higher Temps

Higher Temperatures for Data Centers talks about emerging environmental standards that could well lead to increasing ambient temperatures in facilities. All other things being equal, higher ambient temperatures will lead to higher component temperatures.

In many cases, the maximum processor case temperature (Tcase) is the limiting factor for how high the ambient temperature can be raised. The Tcase limit is established by the semiconductor vendor as the maximum case temperature that the chip can experience and still meet the vendor's reliability goals.

This can put a crimp in plans to use outside air for cooling. In most likely data center locations, there are occasionally warm days that would increase the inlet temperature to the servers to the point that the processor Tcase would exceed the vendor's specified ratings.

The telecommunications market has had this issue for years.
NEBS-rated equipment for central offices generally has to operate at 40°C ambient temperature, but they also need to operate at 55°C for short periods (up to 96 hours at a time and up to 360 hours per year).

To address the needs of the NEBS market, Intel provides some of their processors with
dual Tcase ratings: one long-term T-case rating and a second short-term Tcase rating that is 15°C higher for up to 360 hours per year.

These processors with dual Tcase ratings may be a good fit for systems in data centers that use air-side economizers.

Tuesday, February 3, 2009

Using Outside Air for Data Centers

Conventional wisdom holds that it takes one Watt of cooling to remove every Watt of ICT equipment inside the data center. Today's data centers can do a bit better than that, but cooling costs remain a considerable OpEx cost for facilities.

One approach that's generating increased interest is the use of dry-side economizers, which bring in outside air to cool the data center. Using outside air saves the power that is normally used by compressors and chiller plants to cool facilities; even bigger gains may be achieved by avoiding CapEx (Capital Expense) costs by eliminating the purchase of chiller plants entirely or at least reducing CapEx costs by installing smaller cooling plants.

At first blush, this approach may seem to only be of marginal value. However, higher density data centers (such as those with blade servers) may have a relatively large temperature increase between inlet and exhaust temperature. Even if the desired inlet temperature is only 75°F, a facility with a 50°F temperature rise would have an exhaust temperature of 125°F--most ambient temperatures are well below this temperature. Bringing in outside air could take less energy than cooling the recycled air--humidity considerations notwithstanding.

To look at the impact of using outside air to cool data center equipment, several data center operators have performed small-scale tests to see how data center equipment is impacted by outside air:

Air economizers look promising, based on these results.

--kb