Monday, May 16, 2011
New ASHRAE Temp/Humidity Guidelines
There are even guidelines for lower humidity level if certain procedures are followed.
--kb
P.S. Thanks to Pasi Vaananen for the heads-up.
Sunday, March 7, 2010
Energy Star for Server Should Require Right-sized Power Supplies
Some vendors may ship servers that only draw 200W at 100% utilization with a power supply that can provide 1200W. Traditional ways of evaluating power supplies measure those power supplies across their full rated capacity. If a power supply is sized appropriately, this makes sense. However, a power supply that is too much higher than the system will see in real life should be de-rated.
For examaple, a 2000W power supply might have good efficiency at 50% and 100% of load, but power efficiency tends to drop off at lower load levels, particularly those below 25% of maximum load. Take that same 2000W power supply and put it in a server drawing a maximum of 200W, and the power supply wout always be operating below 10% load. The normal power supply rating levels are of little value if the realistic power draw is much lower than the rated power draw.
To be fair, the tested configurations of servers don't always represent the highest possible loading: adding extra memory, additional hard drives, and extra PCI Express cards can increase a servers power draw. But having no upper limit leaves too much wiggle room and jeopardizes the integrity of the Energy Star rating method.
One possible solution to work around this is as follows:
- Measure the server power consumption under an acceptable benchmark such as SPECpower_ssj2008. Record 2x the maximum power draw (i.e., at 100% load in the benchmark).
- Look at the rated output power for the power supply or power supplies needed to operate the server in that configuration [ignore redundant power supplies used for reliability purposes]. Record the sum of the power of all the non-redundant power supply output power ratings.
- If the answer in Step 2 is less than or equal to the value from Step 1, no adjustment is needed. Skip Steps 4 and 5.
- If the answer in Step 2 is more than the value in Step 1, plot the efficiency rating of the non-redundant power supplies. Extrapolate the efficiency of the power supply (power supplies) at the value recorded in Step 1. Extrapolate the efficiency at 50% and 25% of the value shown in Step 1. Do the same for any other power supply levels normally required, but rate them as a ratio of the value shown in Step 1.
- Evaluate the efficiency of the system based on the load levels and efficiency determined in Step 4 above.
This adjustment would correct ratings for power supplies oversized for the systems they're being tested with. This will incent server vendors to right-size power supplies to better match the real power range of the systems they're being rated for.
--kb
Saturday, February 20, 2010
FaceBook's HipHop Software Efficiency
This showcases two things in particular:
- Software can have a major impact on system efficiency. Even relatively good solutions like PHP can still be improved.
- Metrics that look only at hardware-centric criteria often ignore the benefits of more efficient software.
This second bullet merits further elaboration. Administrators looking at CPU utilization as an approximation of total server work accomplished would erroneously assume their servers were only doing half as much work with HipHop than they were beforehand, even though they would be doing the same amount of work with better software, just doing it more efficiently.
Future posts will talk about ways to measure useful work.
--kb
Tuesday, February 16, 2010
Good IBM doc on cpufreq
Monday, February 15, 2010
Intel® Energy Checker SDK Released
Most of the technology world's focus regarding energy efficiency has focused on hardware: better processors, better memory, better disks, better power conversion, etc. This is good, but it overlooks the substantial contribution that better software can make towards improving energy efficiency. An automobile driver who drives over the top of a hill may use more energy than someone who drives around the hill; software designed with energy efficiency in mind may use a different algorithm than a brute force approach that seems simpler at first.
The Intel® Energy Checker SDK provides developers and systems integrators a simple API that they can use to measure the amount of "useful work" performed by the system and then correlate the useful work with energy consumption. The useful work is not the number of instructions executed, cycles retired, or the average CPU utilization--that's not why you buy software. For example, you buy e-mail software to do things like send e-mails, so the measures of useful work can be the number of messages sent, the number of kilobytes in those messages sent, the number of messages received, and the number of kilobytes in those messages received. Software developers can choose what measures of useful work they export and how often they choose to export this information.
The SDK includes tools to measure the rate of power usage and to measure/calculate energy consumption over time. The SDK supports several external power meters as well as the ability to read energy consumption directly from power supplies having certain levels of instrumentation.
The software developer can easily aggregate/weight the work done in their application(s) with work done in other instrumented applications and compare that to the energy consumed by the system or systems under test to determine energy efficiency. This is an important step towards making software more energy efficient and may lead towards energy-aware algorithms in leading software packages. In turn, this will help administrators measure the aggregate useful work of their facilities, rather than simply measuring hardware-centric metrics that actually penalize more efficient software.
The SDK is available free of charge (and without royalties) from http://software.intel.com/en-us/articles/intel-energy-checker-sdk/. The SDK supports Windows, Linux, Solaris 10, and MacOS X. Source code for the core API and many utilities is included, though Intel distributes some utilities in binary form only. Check it out!
--kb
Thursday, September 3, 2009
Cisco & Sun Servers Spar for Best Humidity Support
Among the major blade vendors, Cisco appeared to have taken the lead by offering support for the broadest operating humidity range, but Sun appears to have matched Cisco recently:
- Cisco's Unified Computing System blades support 5 to 93% RH
- Sun's Sun Fire X2270 blades support "up to" 93% RH
- HP's c-Class BL260 G5 blades support 10 to 85% RH
- IBM's BladeCenter HS22 blades support 8 to 80% RH
- Dell's PowerEdge M600 blades support 8 to 80% RH
(All humidity ranges are non-condensing. All data is from vendor web sites as of 9/3/09).
Ever-widening ranges for supported humidity make the use of dry-side economizers more feasible. If vendors were able to support 0-100% relative humidity, data center operators wouldn't need to worry about humidifcation/de-humidification controls. Eliminating such controls and systems could lower capital costs, reduce operating costs, lower the carbon footprint of facilities, and lower their water footprint as well.
--kb
Thursday, July 16, 2009
Adding a Geographic Element to PUE Calculations
Is a PUE of 1.5 in Phoenix better or worse than a PUE of 1.4 in Seattle?
It depends. In absolute numbers, the lower PUE provides an indicator of the most efficient facility. However, achieving a PUE of 1.5 in Phoenix is much more difficult than an equivalent or slightly lower number in Seattle because Phoenix is so much hotter and requires more air conditioning. Moving data centers to cooler locations helps the PUE rating, but sometimes data centers need to be located in a specific city or region. How can you compare PUE values in regions with different environmental conditions?
One possible approach is to add a geographic compensating factor:
gPUE = G * PUE
The geographic compensating factor G would be determined by The Green Grid or other trusted body based on compiled weather data. Ideally, this could be calculated empirically through a formula using data maintained by the U.S. Department of Energy (refer to this blog link for information on that data and a free tool to visually represent that data).
This approach would allow somebody to measure the technical innovation of a given facility while providing an adjustment to account for geographic disparities in temperature, wind, solar loading, etc. It's not a perfect solution (since some cooling optimizations might not work in cooler or hotter climates), but it provides some measure of equalization to facilitate more equitable comparisons between PUE claims in different locations.
--kb
Monday, June 15, 2009
Making Ice to Lower PUE and TCO
Although some facility managers have developed their own methods for time-shifting HVAC energy requirements, Ice Energy may be the first vendor to market a product specifically designed to do this. The Ice Bear* distributed energy storage system provides up to 5 tons of cooling load during peak hours.
It's good to see innovative products like this coming to market.
--kb
Wednesday, May 27, 2009
Truckin' Down the Information Superhighway
Coincidentally, two days later Amazon introduced Amazon Web Services Import/Export with a blog that starts off with the following colorful quote attributed to Andy Tanenbaum:
Never underestimate the bandwidth of a station wagon full of tapes hurtling down the highway.
Amazon Web Services Import/Export allows people to send USB or eSATA hard drives/media to Amazon for data sets that are impractical to send over available communications links.
It turns out that the bulk version of sneakernet may be the most expeditious way to move data. The more things change, the more things stay the same.
--kb
Note: Revised title on 5/29/09.
Monday, May 11, 2009
How a Good Metric Could Drive Bad Behaviors
Introduction
PUE is defined as follows:
Using the PUE metric, a facility manager can judge what ratio of power is lost in "overhead" (infrastructure) to operate the facility. A PUE of 1.6 to 2.0 is typical, but facility managers are striving to approach a PUE of 1.0, the idealized state.
Companies willing to drive more sustainable practices may incent facility managers to improve facility PUE levels. However, if this is done without context towards the overall energy or other resource consumption, it could drive inefficient behaviors.
Issue #1: Dissimilar Infrastructure Power Scaling
If a facility manager tracks PUE over a variety of workloads, they will see how the data center's infrastructure power consumption tracks with the IT load. Ideally, the infrastructure overhead (HVAC system, UPS system, etc.) will match linearly with the consumption of the servers and other gear in the data center, but this is rarely the case.
In many cases, the fixed overhead for power and cooling systems will become a higher percentage of overall power consumption as the IT load diminished. In other cases, there will be
significant step functions in overall power consumption as large infrastructure items such as chillers, CRACs, or other equipment is turned on or off (as depicted in the graph to the left).In such situations, reducing the IT power consumption could increase the PUE even if it reduces the overall energy consumption of the data center. People will often act in the direction towards which they are incented (i.e., what improves their paycheck). Managers incented to improve PUE without any clear tie-in to overall energy consumption might be reluctant to shut off unused servers or aggressively implement power saving features on their IT infrastructure if it increased their PUE--even if doing so would reduce overall facility power consumption.
Ensuring overall energy consumption is part of the incentive package (not just PUE) is critical to driving the desired behaviors.
[Part of this needs to be linked with overall productivity of the data center so that increased use of the data centers is encouraged while still incenting improved efficiency. I'll write about this in an upcoming post.]
Issue #2: Shifting Infrastructure Loads to IT
Another issue to watch is a desire to classify some infrastructure-like services as IT loads in order to improve PUE efficiencies. Examples of this include moving UPS systems into IT racks or putting large air-mover devices into equipment cabinets and trying to classify them as IT loads. This is "gaming" the system and should be actively discouraged.
The Green Grid is aware of this issue and is adding more guidelines to help people improve the accuracy and consistency of their PUE reporting.
Issue #3: Improving Infrastructure Efficiency at the Expense of IT
The third issue to watch is a move towards facility or equipment practices that reduce the infrastructure power consumption but increase the IT power consumption. In particular, the adoption of higher operating temperatures for data centers warrants particular scrutiny.
I've noted previously that there are significant gains possible by raising data center temperatures and making greater use of dry-side or wet-side economizers. However, it's important to compare the energy savings on the infrastructure side with the energy costs on the IT side. At higher temperatures, leakage currents in silicon increase and fans inside servers need to run faster to move more air through each server.
Increase the IT consumption and lower the infrastructure consumption and you get a two-fer: the PUE numerator goes down and the PUE denominator goes up, lowering the overall PUE. However, if the net power consumption doesn't go down, it usually** doesn't make sense to increase the ambient temperature. Once again, looking at overall power consumption in addition to PUE is important in incenting the proper behaviors.
--kb
**Note: For greenfield (new) data centers or substantial datacenter retrofits, raising the allowed data center temperature may eliminate or substantially reduce the CapEx (capital expenditure) cost for that data center even if the direct energy costs are slightly higher. For example, if a data center doesn't need to purchase a chiller unit, that could shave millions of dollars off the construction cost for a facility. In such cases, more complicated parameters will be needed to evaluate the benefits of raising the ambient temperature in the facility; these likely will include a net present value analysis for the CapEx savings vs. OpEx (operating expense) costs, consideration of real estate savings, etc. The real win is when both CapEx costs are avoided AND OpEx costs are lower.
Wednesday, May 6, 2009
SSD Potential Power Savings Writ Large
However, since the analysis assumed an SSD averaged 7W, SSDs that use less than 2W could save more than 3x that amount. Adding in HVAC and power infrastructure savings, the savings could be even higher.
--kb
Saturday, May 2, 2009
Building Codes and Roof Anchors
If "rooftop renewables" are designed into a building during initial construction, the cost is substantially lower. However, it may not be feasible to install these rooftop renewables when the building is first built (due to limited capital or other reasons).
A middle ground is to provide rooftop anchors during initial construction, regardless of whether or not rooftop renewables are installed with initial construction. That way, solar or wind devices could be installed at a later date much more easily and with no need to breach the roof seal.
All new data centers should be designed for the later installation of rooftop renewables, even if they aren't part of the initial build-out.
Taking this a step further, I advocate the following: building codes should be revised to **REQUIRE** all new commercial buildings with a roof area greater than 1000 square feet to install roof anchors every x (20?) feet, with a TBD load rating for each anchor. (These anchors should also be required when major roof renovations are initiated as well.) Requiring these anchors will facilitate the broader adoption of rooftop renewables in data centers and other commercial buildings.
I hope others will adopt this cause; together we can effect real changes,
--kb
Wednesday, April 29, 2009
Human Side of Higher Data Center Temperatures
One approach is to adopt a "fail in place" model where technicians never go into a production facility, but even Google has technicians adding and replacing individual servers in their containerized data centers.
Other approaches to consider:
- Localized spot cooling. A very small air conditioner could take the edge off the area in front of a rack.
- Perform service operations at night or when it's reasonably cool.
This last suggestion may seem too simplistic at first, but it's actually quite practical. In a facility with sufficient redundancy to ensure high availability, server replacement should be able to wait up to 24 hours. Operating a data center at consistently high temperatures will end up increasing power consumption in the IT equipment. It only makes sense to use higher temperatures in a data center when using optimizers to eliminate or substantially reduce HVAC CapEx and OpEx costs.
If a data center is using economizers, the temperature in the data center should drop when the outside temperature drops. Even in relatively warm areas during summer months, there are substantial times each day where the temperature drops to reasonable levels in which technicians can comfortably work.
--kb
Monday, April 13, 2009
NEBS vs. the Hottest Place on Earth
the hottest place on earth over recorded history was Al Aziziyah Libya in 1922 where 136F (58C) was indicated
James went on to note during his talk that telecommunications equipment designed to the NEBS (Network Equipment Building System) standards routinely has to handle temperatures up to 40°C.
Actually, the story is better than that. NEBS-GR-63 (the key NEBS specification dealing with environmental conditions for equipment in telecommunications central offices) requires equipment to handle 40°C long-term ambient temperatures, but telecommunications equipment certified at the shelf (chassis) level needs to be able to operate at 55°C ambient for up to 96 hours at a time and up to 360 hours per year [the 360 hours is for reliability calculations]. This means that much of the NEBS-rated equipment for data centers can operate at temperatures that are only 3°C lower than the highest natural temperature ever recorded on Earth, as noted by James.
Given the common engineering penchant to provide some guardband on products vs. the official specifications, even a 58°C ambient is not out of the question. This means that NEBS-rated equipment could be good candidates for data centers operating at high temperatures.
But can you get decent performance in NEBS-rated servers? Yes! For example, vendors such as Radisys, Kontron, and Emerson have announced blade servers with Intel's new 5500 (aka "Nehalem") processors, and their bladed servers commonly are NEBS certified to operate at 55°C. This would allow the latest server technology to operate in the most demanding environments.
--kb
Thursday, April 9, 2009
More on Google's Battery-backed Servers
One of the drawbacks not discussed in the prior post is a set of issues related to power transients and harmonics. With a conventional data center, there are multiple levels of power transformation and isolation between the individual server and the grid. Power usually comes in at high- or medium-voltage to a transformer and comes out as low voltage (<600v) before going to a UPS and a PDU.
In an effort to improve efficiency and reduce capital costs, facility managers are looking at removing some of these isolation layers. This is fine to a certain extent. After all, there are a lot of small businesses that run one or two servers on their own, and there aren't major problems with them. In those cases, however, there are usually relatively few computers hooked together on the same side of the electrical transformer that provides power to the building. This transformer provides isolation from building to building (or zone to zone in some installations).
When you scale up into a large data center, however, you get thousands and thousands of servers in the same building. If you remove those extra layers of isolation, the burden for providing that extra isolation falls to the power supplies in the individual servers. If servers use traditional AC power supplies, issues like phase balancing and power factor correction of all the separate power supplies becomes more of an interdepent issue.
The issues can be helped or hurt depending on what's nearby. Servers without isolation near an aluminum smelter, sawmill, subway, or steel mill may see wide fluctuations in their power quality which can result in unexplained errors.
I've seen cases with marginal power feeds where individual racks of servers seem to work fine, but the aggregate load when all servers are operating causes enough of a voltage sag that some servers occasionally don't work right. Let me tell you, those are a real pain to diagnose.
On the other hand, if you're somebody like Google or Microsoft who can locate data centers in places like The Dalles, Oregon or Quincy, Washington that are just a stone's throw from major hydroelectric dams or other sources of power, perhaps you can rely on nice clean power all the time.
External power factors may be the least of a data center manager's problems, however. The big concern with eliminating the intermediate isolation is that transients and other power line problems from one power supply can affect the operation of adjacent systems, and this can build up to significant levels if fault isolation and filtering is not supported.
Another issue that bedevils data center managers is the issue with phase balancing. In most AC-powered systems, power is delivered via three phases or legs (A, B, and C phases), each 120° out of phase with each other. At some point (usually the PDU), a neutral conductor is synthesized so that single-phase currents can run from one of these legs to neutral. In a properly balanced system, there will be equal loading on the A leg, the B leg, and the C leg. If the phases are not properly balanced, there are several bad things that can occur, including the following:
- The neutral point will shift towards the heaviest load, lowering the voltage to the equipment on that line, resulting in premature equipment failure and undervoltage-related errors
- An imbalanced load may cause excess current to flow over specific conductors and overheat
- Breakers or other overcurrent mechanisms may trip
Phase imbalance can occur when network administrators do not follow a rigorous process of plugging every third server into alternate phases. Additionally, shifting workloads could cause some servers to be more heavily utilized than others--and phase balancing is almost certainly not a factor considered in allocating applications to specific servers. An even more pernicious issue can arise with systems employing redundant power supplies, such as blade servers: in an attempt to maximize efficiency, management software may shut down certain power supplies to maximize load on the remaining power supplies--all without considering what the impact to phase balancing is when the load is not equally shared among all power supplies.
Data centers that employ conventional PDUs don't generally have these issues (or have them at lesser severity), since the PDUs and their transformers are usually designed to handle significant phase imbalances without creating problems.
Additional considerations with the Google battery-backed server approach:
- Acid risks from thousands of individual tiny batteries (i.e., cracked cases in thinner-walled batteries)
- Shorting risks from batteries that can deliver thousands of amps of current for a short period
- More items to monitor, or higher risks of silent failures (albeit with smaller failure domains) when you most need the batteries
This is a complex issue. I'm not convinced that Google has determined the optimal solution, but kudos to them for finally being willing to publicly discuss some of what they consider to be best practices. Collectively, we can learn bits and pieces from different sources that could end up delivering more efficient services.
--kb
Saturday, April 4, 2009
Evaluating Google's Battery-backed Server Approach
Are batteries in servers a good idea?
There are some definite advantages in Google's approach:
- No need to pay for UPS systems (saves CapEx dollars)
- Eliminates two conversion stages found in a traditional AC double-conversion UPS
- Reduces dedicated floor space/real estate commonly devoted to UPS/battery rooms
- Localizes fault domains for a failed server to just one server
- Scales linearly with the number of servers deployed
All of these add up to a solution that works just as well for one server as it does for one thousand servers. Coupled with Google's efforts to increase energy efficiency through founding and support for the Climate Savers Computing Initiative (CSCI) and its target of 92% power supply efficiency, this solution appears to be very efficient.
However, there are some down sides to Google's approach:
- A lot of batteries to wire up and monitor
- Increased air impedance from blocking airflow
- Lower battery reliability with increased ambient temperatures
- Higher environmental impact due to increased battery materials
- Individual server supplies are exposed to a higher level of power transients and harmonics
- Potential phase imbalances and stranded power in data centers
Issue #1 is self-obvious. Issue #2 can be seen from this picture from Green Data Center Blog; the physical mass of the batteries blocks a good portion of the air space in front of the server, which increases the resistance and in turn requires more fan power to move the same amount of air.
Issues #3 and #4 are somewhat related. Google, Microsoft, and other leading internet companies have advocated moving the ambient temperatures of data centers to higher temperatures, with some advising 35°C, 40°C, or even occasionally 50°C ambient temperatures. There are clear savings to be had here, but it may run counter to the battery approach used by Google. Assuming the Google batteries are conventional lead-acid batteries, a common rule is that the useful life of batteries drops by ~50% for every 10°C above 25°C ambient temperatures. Thus, a 4-year battery would only be good for ~2 years in a 35°C environment. In comparison, conventional UPS batteries are often rated for 10, 15, or 20 years. When consolidated in a UPS battery cabinet, the batteries can be protected from the higher ambient temperatures through localized cooling (batteries dissipate almost no heat) for increased life.
Lots of little batteries like Google uses results in more materials usage compared to the use of larger batteries. Couple that with reduced battery life at higher temperatures, and the result is not as good as it first seems. According to http://www.batterycouncil.org/LeadAcidBatteries/BatteryRecycling/tabid/71/Default.aspx, more than 97% of lead from lead-acid batteries is recycled, but this also states that 60-80% the lead and plastic of new batteries is recycled material. Looking at this last stat a different way, 20-40% of lead-acid battery materials are not recycled. Thus, even if Google performs 100% battery recycling, using lots of new batteries still results in the use of a lot of new materials.
I'll address issues #5 and #6 in a future post.
--kb
Friday, April 3, 2009
Google's Server Power Supplies

One of the more interesting aspects revealed Wednesday was the fact that Google has batteries attached to each of their servers.
At first, this seems rather odd. Google's explanation for this is that they use this arrangement as a 99.9% efficient replacement for UPS (Uninterruptible Power Supply) systems. Wow...99.9% efficient!
This is definitely a different approach from what most data centers do today, and it seems really far out there--until you break it down in its component parts. A simplified block diagram looks like the following:

- External power supply provides ~12Vdc
- Battery is included with every computer
- When the external power supply fails, the battery provides power until the generator starts or power is switched to a different source
Graceful shutdown in power outages may or may not be an issue for Google's applications (likely not an issue).
Wednesday, April 1, 2009
Deciphering Intel Code Names
Okay, but how can you find out what each one of these code names refers to? Well, it turns out that Intel has a web site that allows you to enter code names for released products and then look up the relevant information. Go to http://ark.intel.com/ and enter the code name (or official name) of a current Intel product, and chances are it will be listed.
One particularly useful feature of this site is the System Design capability. For example, if you enter a processor/chipset power budget and other criteria, the site will list all matching combinations. Try it out!
--kb
Tuesday, March 31, 2009
Suggestion for Energy Star Measurement of Blade Power Consumption
Ideally, there would be a standardized benchmark like SPECpower_ssj2008 that would be able to measure power consumption on a per-blade basis, but the current benchmark doesn't have provisions to handle chassis.
As an alternative, here are suggestions for how the EPA could measure power consumption for Energy Star (until a chassis-friendly industry specification is developed by an industry group like SPEC):
- Apply Energy Star to blades, not to chassis. Chassis are ineligible to meet Energy Star, but the blades that go in them can be Energy Star certified.
- Configure a chassis with the minimal amount of chassis management modules and external modules required for operation, but include all supported power supplies for a given chassis and all the fan/cooling modules typically used (don't remove redundant fans or power supplies).
- Run a sample workload on all servers to keep them minimally active. Install the same server configuration in all server slots.
Measure total power consumption to all power feeds in the chassis under two conditions and with the following calculations:
- Condition 1: Determine power consumption P1 with all N server blade slots installed.
- Condition 2: Remove servers so that N/2 (round up) servers are evenly distributed in the chassis; call that number N'. Determine power consumption P2 at this level.
- P3 = P1 / N. This is the weighted average power per server blade in a full chassis.
- P4 = P2 / N'. This is the weighted average power per server blade in a half-full chassis.
- P5 = (P3 + P4) / 2. This is the weighted average power per server blade.
Notes:
- This accounts for chassis overhead, including fans, power supplies, management modules, and network connectivity. There is a slight penalty to blades here since rack-mount servers don't include any allocation for network switch power, but represents the minimum configuration needed to use those blades. Additionally, many vendors have low-energy networking elements (i.e., passthrough blades) that minimize this impact.
- If the chassis contains power supplies to convert input voltages to a different voltage supplied on the backplane, the power supplies used in the chassis must meet the power supply qualification requirements outlined elsewhere in the Energy Star for Servers specification.
- If a chassis contains redundant power supplies, the server blades are eligible for an allowance of 20W per redundant power supply, divided by the number of servers. For example, if a chassis has 2+2 power supplies (2 redundant power supplies and 2 minimum power supplies for a fully loaded chassis) and 10 blades, then each server would get a 4W/server allowance (2 * 20W / 10 servers).
With all the notes above, this may look to be complicated, but it's actually a fairly simple configuration that provides a close analog to how standalone rack-mount servers are tested. This could be used in the initial version ("Tier 1") of the Energy Star for Servers specification if the EPA wanted to use it.
--kb
Thursday, March 12, 2009
Eliminating the UPS Efficiency Penalty with -48Vdc: Part II
Let's start by looking at the power supply unit (PSU) component by itself. Based on the information in the quantitative analysis by The Green Grid, high-efficiency AC and DC power supplies look like this when compared to each other:

The graph shifts to the right when redundant power supplies are considered. Since there are numerous different voltage converters in a server (modern servers often have in excess of 25 voltage rails used internally), it's really impractical to try to duplicate every voltage converter in a server--at least if you want it for a reasonable price. However, servers with redundant power supplies provide three principal benefits:
- Connectivity to separate primary power sources (i.e., different utility feeds)
- Protection against failure in upstream power equipment (i.e., failure in a PDU)
- Cabling problem or service failure (i.e., accidentally unplugging the wrong server)
In contrast, a DC system has no phasing issues to deal with. Therefore, DC-based equipment has two main options: full duplicate power supplies (like AC) or using a technique called diode OR'ing (or FET OR'ing) to safely combine power from two separate DC sources as inputs to a single power supply. [Since there are numerous downstream power converters that are not redundant, there's no need for the power supply itself to be redundant--it just needs to be fed from multiple inputs.] Many DC power supplies do this today, as this approach is commonly used in the highly-reliable telecommunications system with -48Vdc systems. The result is a wider gap between the net AC power supply efficiency and the DC power supply efficiency:

Taking this a step further, look at the typical operating point for servers vs. their power supply ratings. For example, look at the various published reports for SPECpower_ssj2008: you'll notice there are numerous cases where the power supply shipped with the system is 2-4 times the maximum power draw in the sytem. If the power supply in a system is 2x the necessary power, then the system would normally operate in the left half of the graph immediately above. If the average power is considerably less than the maximum power draw, then the system could spend the bulk of its time operating at the 25% load level or less in the graph above.
At these lower loads, the efficiency benefits of -48Vdc systems become more apparent, even when there's no UPS in the picture. If an installation uses UPSes, the efficiency gap widens further in favor of -48Vdc.