- IT equipment turns nearly every watt it draws into heat, so a cooling failure in a small room can push temperatures to a damaging level in minutes, not hours.
- The right system type depends on load density and room size: a ductless split can carry a small closet, while a dense row of racks usually needs in-row or computer-room air handlers.
- Airflow management (blanking panels, sealed floor openings, hot- or cold-aisle containment) often recovers more usable capacity than adding another cooling unit.
- Redundancy is a design decision made before the failure: N+1 capacity, separate electrical feeds, and alarms that reach a person are what keep a single fault from becoming an outage.
- Monitoring should be independent of the equipment it watches. Temperature probes at rack inlets, leak detection, and an alarm path that works when the network is down matter more than a pretty dashboard.
- Maintenance on precision cooling is mostly about airflow, condensate, filters, belts, and controls. Most mission-critical trips start as a small, visible problem someone did not log.
Why server room cooling is a different problem from comfort cooling
Comfort air conditioning is designed around people: it removes the heat from occupants, lighting, sun through glass, and outside air. The load is moderate, it swings with the weather, and nobody is harmed if the space drifts a few degrees for an afternoon. A server room inverts every one of those assumptions. The load is steady and almost entirely sensible heat, meaning it raises air temperature without adding moisture. It runs all day, every day, and it is concentrated into a small footprint.
Almost every watt delivered to a server, switch, or storage array leaves as heat. A rack drawing a modest amount of power releases roughly the same amount of heat into the room, and the cooling system has to move that heat out continuously. Because the heat arrives at a constant rate regardless of the weather, the cooling plant cannot simply be sized from square footage the way an office is. It is sized from the electrical load of the equipment, plus growth, plus the redundancy you decided to carry.
The second difference is time. In a conventional office, a failed rooftop unit is an uncomfortable afternoon. In a closet with a handful of racks and no airflow path to the rest of the building, inlet temperatures can climb rapidly once cooling stops, because there is very little air mass to absorb the heat. Many servers throttle or shut themselves down to protect hardware, which turns a mechanical problem into a business outage. That is why the sections below spend as much time on detection, redundancy, and response as on the equipment itself.
This guide is general background for facility and IT staff. It is not a design document. Operating temperature and humidity ranges, redundancy tiers, and fire-protection interlocks should come from the equipment manufacturers, your IT standards, and the engineer of record. Confirm current requirements for your site.
The main cooling system types
Precision cooling equipment falls into a handful of families. Each solves a different density and budget-of-space problem, and many facilities end up with a mix because the room grew in stages.
| System | How it works | Best fit | Watch for |
|---|---|---|---|
| Ductless or commercial split (wall, ceiling cassette, floor) | Indoor unit moves room air over a refrigerant coil; outdoor condenser rejects heat | Small server rooms, telecom closets, IDF/MDF spaces, edge sites | Condensate removal, low-ambient operation, no built-in redundancy unless two units are installed |
| CRAC (computer room air conditioner) | Self-contained refrigerant unit with its own compressor, fans, and humidity control, usually floor-mounted with a remote condenser | Rooms with raised floors or ducted supply, moderate loads | Compressor wear, refrigerant leaks, reheat and humidifier maintenance |
| CRAH (computer room air handler) | Fan and coil unit supplied with chilled water from a central plant or dedicated chiller | Larger rooms and buildings that already have a chilled-water loop | Dependence on the plant, valve and sensor calibration, water leak risk near equipment |
| In-row cooling | Compact cooling units sit between racks and draw hot air from one side, supplying cold air to the other | High-density rows, retrofits with no raised floor | Piping or refrigerant routing in the white space, condensate, coordination with containment |
| Rear-door heat exchangers | Water-cooled coil mounted on the back of a rack captures exhaust heat | Isolated hot racks in an otherwise ordinary room | Water near electronics, flow monitoring, dew point control |
| Packaged rooftop or split with economizer | Standard commercial equipment adapted for a dedicated space | Light-duty rooms where precision control is not required | Humidity swings, dust and outdoor air quality if free cooling is used |
A small room with a few racks is usually well served by a pair of commercial ductless or split systems, set up so either can carry the load alone. Move up to several racks with real density and you start to need dedicated precision units that can hold a narrow temperature band and manage humidity as part of the same control loop. Chilled-water CRAH units make sense when a central plant already exists and has the capacity, a situation covered in more detail in our chillers and hydronic systems guide.
One distinction matters in every system: sensible versus total capacity. Comfort equipment spends a significant share of its capacity removing moisture. A server room has almost no latent load, so precision units are rated by sensible capacity, and a comfort unit that looks adequate on paper can fall short once you compare the number that actually matters.
Airflow management and containment
Before buying more cooling, look at what the existing cooling is doing. In many rooms, a large fraction of the cold air never reaches a server inlet. It leaks through cable cutouts, bypasses racks through open slots, or mixes with hot exhaust before it is ever used. The visible symptom is a room that feels cold near the unit and still has hot spots at the tops of racks.
The usual remedies are inexpensive compared with new equipment and should be done first:
- Blanking panels in every unused rack position, so exhaust air cannot recirculate to the front of the cabinet.
- Sealing floor and wall penetrations and cable cutouts, so supply air is not lost to places that do not need it.
- Hot-aisle/cold-aisle layout, with servers drawing from the cold aisle and exhausting into the hot one, rather than a mixed orientation.
- Containment of either the cold aisle or the hot aisle with doors, roofs, and end caps, which keeps the two air streams from mixing at all.
- Return path discipline, so hot air has a direct path back to the cooling unit coil instead of wandering across the room.
Containment lets supply temperatures be raised without creating hot spots, which improves the efficiency of the cooling equipment and often reduces compressor run time. It also changes how you test. Measure at the server inlet, not at the unit return. A return air reading in a contained system can look hot because it should be, while the inlet reading is what governs equipment safety.
Containment interacts with fire protection. Roofs and doors on an aisle can affect detector coverage and suppression discharge, and local fire authorities may have requirements for them. Treat this as a coordination item with the fire protection contractor and the authority having jurisdiction, and confirm current requirements before installation.
Temperature and humidity: thinking in terms of an envelope
Equipment makers publish an operating envelope for inlet temperature and humidity, and industry bodies such as ASHRAE publish recommended and allowable ranges for IT environments. The details change with equipment generation and class, so this guide does not quote a number. The practical points are these.
- The envelope is measured at the server inlet, not at the thermostat on the wall.
- Operating slightly warmer than old habit suggests is common in modern designs, but only with good airflow management. Warmer supply air in a leaky room just creates hot spots.
- Humidity matters at both ends. Very dry air raises static discharge risk, while high humidity and low surface temperatures create condensation risk on cold pipes and chilled-water components.
- Rate of change matters too. Equipment is more tolerant of a stable reading than of rapid swings caused by short-cycling or a poorly tuned control loop.
- Dust and corrosive contaminants count. Rooms next to kitchens, loading docks, or coastal air intakes can accumulate particulates or salts that shorten electronics life.
Humidity control is often the most neglected part of the system. Humidifiers (steam canister, infrared, or ultrasonic) and reheat elements are service items with water quality, scaling, and electrical considerations. In many small rooms, tight vapor barriers and a sealed envelope mean humidity does not need aggressive control, and a competing set of humidifier and dehumidifier cycles in two nearby units wastes energy and causes drift. If you run more than one unit in a room, make sure their setpoints and control modes are coordinated rather than fighting each other.
Redundancy: designing for the failure you will eventually have
Every mechanical system eventually fails, and the question is whether the room survives while it is repaired. Redundancy is the design response, and it lives at several levels at once.
| Approach | Meaning | Typical use |
|---|---|---|
| N | Cooling capacity equals the load, with no spare | Non-critical rooms, short-term situations |
| N+1 | One extra unit beyond what the load needs, so any one can fail | Most business-critical server rooms |
| 2N | Two fully independent systems, each able to carry the load | Critical facilities with strict uptime goals |
| Distributed | Many small units sharing the load with automatic rotation and lead-lag | Edge sites and closets |
Redundant units only help if they are actually independent. Two cooling units on the same electrical panel, the same condensate pump, or the same chilled-water branch share a failure mode. Review the electrical distribution, the control power, the condensate path, and the monitoring for single points of failure, not just the nameplate count of units.
Rotation matters as well. If one unit always runs and the other is a cold spare, the spare may fail to start when called, because nobody has run it in a year. Lead-lag controls that alternate duty, plus a routine test where a unit is deliberately shut down while the other proves it can carry the load, turn a theoretical spare into a verified one.
Utility power is the other half of this. Cooling needs to ride through the same events that the IT equipment rides through. If the servers are on a UPS and generator, but the cooling units are not, the room may hold power and lose cooling. Plan the transfer sequence, restart behavior after power returns, and whether chilled-water pumps and controls are on emergency power. These are electrical coordination items to settle with the electrical engineer.
Monitoring and alarms that actually reach someone
A cooling failure at three in the morning is only survivable if someone learns about it in time to act. Monitoring has three jobs: tell you what is happening now, warn you before limits are reached, and reach a human who can respond.
- Rack inlet temperature sensors, ideally at the top, middle, and bottom of representative racks, because the top is usually the hottest.
- Supply and return temperatures at each cooling unit to see capacity trends.
- Humidity at a few representative points.
- Leak detection near chilled-water piping, humidifiers, condensate pumps, and under raised floors.
- Unit status from the cooling equipment itself: compressor running, fan status, alarm codes, filter differential pressure.
- Power status for cooling units and their supporting pumps.
Design the alarm path to survive the conditions it is warning about. If alerts depend on the same network switch that sits in the overheating room, the notification may die with the room. Use independent monitoring where possible, with cellular or separate-path notification and an escalation list that includes more than one person. Test it on a schedule by actually triggering an alarm and confirming that someone receives it.
Alarm thresholds need two tiers. An early warning, set well inside the equipment envelope, gives people time to investigate. A critical alarm, set closer to the limit, triggers the response plan. If every alarm is critical, people learn to ignore them. If thresholds are set too loose, the first alarm arrives when there is no time left.
The DaVinci Portal can store the history behind a cooling unit: service records, readings, photos, and the asset record linked from a QR tag on the equipment. That does not replace live environmental monitoring, but it gives the technician arriving after an alarm the recent history without a phone call.
Failure timelines: what happens when cooling stops
How long a room can survive without cooling depends on load density, room volume, air leakage, adjacent spaces, and starting temperature. A large room with a modest load behaves very differently from a small closet with dense racks. Instead of quoting minutes, the useful exercise is to measure your own room.
- Document the load. Total electrical draw for the room, from UPS or PDU readings, gives the heat input.
- Run a controlled test where safe. With IT leadership present and a plan to restore cooling, record how inlet temperatures rise over a short interval after shutting down one unit. Extrapolating from a partial test beats guessing.
- Identify the first thing to give. Often it is not the servers but the UPS batteries or a storage array with a lower threshold.
- Write the response ladder. Open doors, add portable cooling, shed noncritical load, graceful shutdown of lower-priority systems, then full shutdown, each with a trigger temperature and an owner.
Common mechanical causes of sudden loss include a tripped breaker or control fuse, a failed contactor, a frozen evaporator coil from airflow or refrigerant issues, a high-pressure lockout from a dirty or blocked condenser, a clogged condensate pump that trips a float switch, and, on chilled-water systems, a pump failure, a closed valve, or loss of the plant. Slower failures are quieter: a gradual loss of refrigerant charge, fouled coils, and failing fan bearings reduce capacity for weeks before anyone notices.
Keep a plan for portable cooling. Know where spot coolers could be placed, what power they need, and how the exhaust or condensate would be handled. Confirm access, electrical availability, and rental lead times ahead of time rather than during the event.
What maintenance on precision cooling should cover
Precision cooling is maintained more like a process system than an office unit. The goal is to find drift before it becomes a trip. A good scope combines routine inspections, instrument verification, and operational tests.
| Task | Why it matters |
|---|---|
| Filter inspection and replacement | Restricted airflow lowers capacity and can freeze the coil |
| Coil cleaning, evaporator and condenser | Dirt raises head pressure and cuts heat rejection |
| Condensate drain, pan, and pump check | A blocked drain trips float switches and can leak into the room |
| Fan, belt, and bearing inspection | Vibration and wear show up well before failure |
| Electrical connections, contactors, capacitors | Heat-stressed parts cause most sudden no-cooling calls |
| Refrigerant pressures, superheat, subcooling | Detects slow leaks and metering problems |
| Humidifier and reheat service | Scale, canisters, and heaters drift out of spec |
| Sensor and setpoint verification | A drifting sensor makes a healthy unit act sick |
| Failover and alarm test | Proves the spare and the notification path work |
| Records update | Builds the evidence base for trend analysis and audits |
Schedule work with the IT team. Some tasks need a unit offline, which means the remaining unit carries the room alone. That is the moment to confirm redundancy is real, but it needs a change window, a rollback plan, and a person watching inlet temperatures. For central-plant CRAH systems, coordinate valve and pump work with whoever owns the plant.
A preventive maintenance program built around these tasks, with reports tied to each unit, is the practical way to catch drift. Our maintenance plans are written around the equipment you actually have, not a generic checklist.
Refrigerants, leaks, and the regulatory background
Precision cooling equipment often holds a significant refrigerant charge, and the refrigerant landscape is changing. Federal and California rules under the AIM Act, EPA Section 608, and CARB programs are phasing down high-global-warming-potential refrigerants and setting leak-repair and recordkeeping expectations. The specifics depend on charge size, equipment type, installation date, and refrigerant, and they change, so confirm current requirements with the agencies or your compliance advisor.
- Know what is in each unit. Refrigerant type and charge are on the nameplate, and your asset records should match.
- Log leaks and repairs. Equipment above certain charge thresholds typically triggers leak inspection and repair recordkeeping.
- Plan for the future supply. Older refrigerants face supply constraints as phasedown schedules progress, which affects repair-versus-replace decisions on aging CRAC units.
- Use certified technicians. Handling refrigerant requires EPA certification, and recovery and reclaim records are part of the paper trail.
- Consider safety classification. Newer lower-GWP refrigerants may carry flammability classifications that change room ventilation, detection, and servicing requirements.
Our refrigerant compliance guide for California goes into the recordkeeping side in more depth. For a server room, the key point is practical: a slow leak shows up first as lost capacity and rising compressor temperatures, which is a reliability problem before it is a compliance problem.
Choosing and working with a mechanical contractor for IT spaces
Work in occupied IT spaces has its own etiquette. Technicians need to understand escort rules, anti-static practices, no-touch zones, and the difference between a service call and a change request. A contractor who treats a server room like a mechanical room will eventually bump a cable or trip a breaker.
- Ask how they isolate a unit without losing the room, and how they confirm the remaining unit is carrying the load.
- Ask for a root-cause explanation on every repeat failure, rather than another part swap. Our diagnostics approach is built for exactly this.
- Insist on documentation: readings before and after, photos, and a record per unit.
- Plan change windows in advance, with IT, security, and facilities in the loop.
- Confirm credentials: California contractor license, refrigerant handling certification, and insurance suited to the site.
Davinci Mechanical is a commercial-only contractor, license CA #1083101, with union (UA Local 250) labor and documentation-first service. For sector context, see our pages on data centers and telecom facilities. If an existing cooling problem keeps coming back, a second-opinion diagnostic is a good starting point, and larger replacement or expansion work is handled through our bid and project process.
Frequently asked questions
What is the difference between a CRAC and a CRAH unit?
A CRAC contains its own refrigerant compressor and cools air directly. A CRAH has no compressor and uses chilled water from a central plant or chiller to cool the air. CRAHs depend on that plant, while CRACs are more self-contained.
Can a regular split system cool a server room?
For small, low-density rooms, commercial ductless or split systems are commonly used, ideally in a pair so either can carry the load. Denser or humidity-sensitive rooms usually need precision equipment. The deciding factors are sensible capacity, redundancy, and control.
How warm can a server room run?
It depends on the equipment and the standard you follow. Manufacturers publish operating ranges for server inlet temperature, and industry guidance defines recommended and allowable envelopes. Check your equipment documentation, measure at the inlet, and confirm with your IT standards.
Do we need two cooling units?
If an outage would stop the business, redundancy is the standard answer, typically N+1 or better. The second unit has to be electrically and mechanically independent to count, and it should be tested regularly.
How often should server room cooling be serviced?
Most critical rooms are checked on a quarterly or more frequent cycle, with additional annual tasks such as electrical and refrigerant circuit verification. The right frequency depends on load, redundancy, and how the room is monitored.
What should we do if the cooling fails tonight?
Follow your written response ladder: confirm the failure, check for tripped breakers and alarms, open doors or deploy portable cooling, shed noncritical load, and call your service provider. Having the ladder written, with owners and trigger temperatures, matters more than any single step.
Have a site this applies to?
Davinci Mechanical is the commercial and union division of Scottish Tom's Heating & Air. Send us the equipment list or the problem and we'll tell you what we'd check first.