Projects · Mini Hostpapa

Datacenter & Physical Network Foundations

Colocation, racks, power, cabling, top of rack switching, VLAN trunking, bonding, MTU, and out of band management, with photographs of the actual hardware.

Updated Aug 3, 2026 · 70 min read

Datacenter & Physical Network Foundations

Okay. Docs 1 and 2 were about concepts and identity. Packets, addresses, routing, who owns what.

This one is about metal, copper, fiber, watts, and rack units. It is the part where a mistake costs a drive to the datacenter instead of a config reload.

The rule of physical work

You get one clean install window.

Cables you label badly today, you re trace at 3 in the morning next year with a flashlight held in your teeth. Physical discipline is not aesthetics and it is not perfectionism. It is mean time to repair, and mean time to repair is the number your customers actually experience during an incident.

Volume 0 already showed you the rack at executive level: one 42U cabinet in a Tier III or IV class facility in Tunisia, an edge firewall, a top of rack switch, a small out of band switch, two CloudStack management servers, three or four KVM hosts, and one NFS storage node. That page deliberately said "VLAN IDs come later."

Later is now. Here we turn that sketch into something you could hand to a colo engineer and a hardware vendor without flinching.

How this document is built, and why

I have walked datacenter floors, and the honest truth is that reading about this hardware is nothing like seeing it. So every section here shows you the actual metal, with photographs, and tells you what to look at in each one.

If you have never racked a server, the goal is that by the end of this page you can recognize every component, name it, and explain what it is for. That is most of the gap between someone who has done it and someone who has only read about it.

The two things that surprised me most on a datacenter floor, and nobody writes them down: it is genuinely cold, cold enough that you want a jacket in July, and it is loud, loud enough that you have to raise your voice to be heard next to a rack. Both of those are engineering choices you are about to learn the reasons for.


What you will be able to produce after this document

OutputWho consumes it
Rack elevation diagram with U positions and weight distributionColo provider, install technician
Power budget in kW with A and B feed splitColo contract, finance
Port map and cabling schedule with labelsRemote hands, and future you
VLAN table with IDs, subnets, gateways, and purposeDoc 4, Doc 5, and Volume 5
Switch and firewall hardware requirements listProcurement
Out of band access planSecurity review, on call runbooks

1. Colocation, decoded

The analogy: renting a serviced office

You do not build your own office tower to run a ten person company. You rent a unit in a building where somebody else handles the power, the air conditioning, the lifts, the security desk, and the fire suppression.

You bring the desks, the computers, and the locks on your own door. And when you need something moved while you are away, you pay the building concierge by the hour.

Colocation is exactly that, for servers. The provider runs the building. You own everything inside your cabinet, and you own every decision about what crosses your cabinet boundary.
Loading media
A datacenter aisle between two rows of black server cabinets with perforated mesh doors, standing on a raised floor with perforated floor tiles running down the middle of the aisle.
A colocation floor. Look at three things: the perforated mesh doors, the raised floor, and the perforated tiles running down the middle of the aisle. Every one of those is about moving air. Source: Wikimedia Commons, CC BY-SA 3.0.

Study that photo for a second, because it contains most of section 2 already. The doors are mesh, not solid, so air can pass through. The floor is raised, with removable tiles. And the tiles in the walkway are perforated, while the tiles under the cabinets are solid.

Loading media
Close up of the underside of a perforated raised floor panel showing the grid of ventilation holes and the panel's structural frame.
The underside of a perforated floor tile. Chilled air is pushed into the void beneath the floor and rises through these holes into the aisle where the servers breathe. Source: Wikimedia Commons, CC BY-SA 3.0.
Why the aisle you stand in is the cold one

This is the design that made the "it is genuinely cold in there" observation click for me.

Chilled air is pumped into the void under the raised floor, rises through the perforated tiles into the aisle, gets sucked in through the front of every server, absorbs heat from the components, and is blown out the back into the other aisle.

So a datacenter has cold aisles, where server fronts face each other and you stand, and hot aisles, where server backs face each other. Walk from one to the other and the temperature difference is dramatic and immediate. All servers in a row must face the same way, or you get a server inhaling its neighbour's exhaust, which is exactly as bad as it sounds.

What colo actually sells you

What you buyWhat it means in practiceWhat to check in the contract
SpaceA cabinet, a locked cabinet, a cage, or a private roomCabinet height in U, and depth in mm, because deep servers do not fit shallow cabinets
PowerA committed number of kilowatts per cabinet, on one or two feedskW per rack, whether it is A and B redundant, and what happens if you exceed it
CoolingKeeping the cold aisle at a target temperatureWhether your power density is actually coolable in that row
Physical securityAccess control, cameras, escorts, loggingWho can open your cabinet, and whether that is logged
Cross connectsA physical cable from your cabinet to a carrier or an exchangePrice per cross connect, and the lead time to install one
Remote handsTheir technician physically touching your gear on your instructionHourly rate, minimum billing increment, and response time commitment

The bits of the building you will never own, but must understand

Loading media
Rows of large sealed backup batteries mounted on metal racking inside a datacenter UPS room.
The UPS battery room. These carry the entire building for the minutes between a utility power failure and the generators reaching full load. Source: Wikimedia Commons, CC BY 3.0.

Here is the chain of events on a utility power failure, and understanding it is what lets you read a Tier claim honestly: utility power dies, and the UPS batteries take the load instantly, because there is no time to wait for anything mechanical. The generators start and take perhaps thirty seconds to a minute to stabilize. Once they are stable, the load transfers to generators, and the batteries recharge. The batteries are not there to run the building for hours. They exist to cover the gap until the generators are ready.

Reading Tier claims like an engineer instead of a buyer

Tier is an industry classification for how redundant the power and cooling are. Higher is stricter. You do not need to memorize the standard, and a provider saying "Tier IV" on a brochure is marketing until you read the contract.

What actually matters, and what you should ask on the tour:

Are the two power feeds genuinely independent all the way back, or do they merge somewhere upstream? A single point of failure hidden three floors down is still a single point of failure.

How often are the generators load tested, and can you see the records? A generator that has never been tested under load is a decoration.

Is cooling redundant, and what is the actual failure behaviour? Losing cooling gives you minutes, not hours.

How many carriers are in the building, and can you reach more than one? This is the physical prerequisite for the multihoming discussed in Doc 2. No second carrier in the building means no ASN qualification, no matter what you want.

The meet me room and cross connects

Side definition: meet me room and cross connect

The meet me room is the space in a colocation facility where all the carriers terminate their networks. It is the marketplace of the building.

A cross connect is a physical cable run from your cabinet to somebody else's equipment in that room, usually installed by the provider's technicians and billed monthly.

Why this matters commercially: the number of carriers in the meet me room is the number of transit providers you can realistically reach without expensive external circuits. A building with one carrier is a building where you can never be multihomed. Ask this question before you sign, not after.

The questions to read aloud on the tour

Print these. Genuinely.

QuestionWhat a bad answer sounds like
How many kW per cabinet, and is that per feed or total?Vagueness. This number must be exact, because your budget in section 3 depends on it.
Are A and B feeds independent to the building entrance?"They are both redundant" without explaining the path
Which carriers are in the meet me room today?A list of who they are "in discussions with"
What is the cross connect fee and lead time?No published price
What is the remote hands rate and response commitment?Best effort with no number attached
Who can physically open my cabinet, and is it logged?Anything other than a specific, auditable answer
Can I see generator test records?Hesitation
What is the cabinet depth in millimetres?"Standard." There is no standard depth.

That last one sounds trivial and it is the question that saves an install day. This ties directly into the physical controls we committed to in governance, risk and compliance, where physical security is an audited control rather than an assumption about somebody else's building.


2. The rack as an engineering artifact

The analogy: a bookshelf with standardized slots

A rack is a metal bookshelf where every shelf position is a standard height, so any manufacturer's equipment fits any rack. That standardization is the entire reason a hardware ecosystem exists.

The slot is called a rack unit, written U.

Loading media
Engineering drawing of a 19 inch rack mounting rail showing threaded holes, hole spacing of 15.875 mm and 12.7 mm, and bracket dimensions for 1U at 44.45 mm, 2U at 88.9 mm, 3U at 133.35 mm, and 4U at 177.8 mm.
The actual dimensions. 1U is 44.45 mm, and each U is three holes with uneven spacing. Source: Wikimedia Commons, CC BY-SA 4.0.

At first I found the 1U measurement genuinely strange, and here is the trick to understanding it: 44.45 mm looks like a random number until you realize it is exactly 1.75 inches. The whole standard is imperial, defined decades ago, and the metric numbers are just conversions. That is why nothing about it is round.

Now look more carefully at that drawing, because it contains a detail that trips up every first timer. Each U has three mounting holes, and they are not evenly spaced. The gaps go 15.875, 15.875, then 12.7 mm before the next U begins. That uneven gap is how you can tell where one U ends and the next starts.

The first timer mistake this causes

Because the holes are unevenly spaced, it is entirely possible to mount equipment straddling a U boundary, using the wrong three holes. Everything appears to fit. Then the next device up will not mount, because the remaining holes no longer line up with a U boundary, and you have to unrack and redo it.

Good cabinets have U numbers printed on the rails. Use them. Count from the bottom, and mount to the printed numbers rather than by eye.

Height is the easy dimension. Depth and rails are the hard ones.

DimensionWhat to knowWhat goes wrong
Width19 inches between rails, effectively universalRarely a problem
HeightMeasured in U. A full cabinet is typically 42U or 47U.Running out of U, which is a planning failure not a surprise
DepthVaries significantly. Deep servers need roughly 1000 mm or more of cabinet depth.The server physically does not fit, or fits with no room for cable bend radius behind it
Rail typeSquare hole, round threaded, or tapped. Your rail kits must match.You arrive with rail kits that do not fit the cabinet and can do nothing about it that day

Depth is the dimension that ruins install days because everybody remembers to count U and nobody thinks to measure how deep the cabinet is. Ask for cabinet depth in millimetres and compare it against the deepest server on your purchase order, then add room behind it for cables to bend.

Weight, and why the heavy things go at the bottom

Loading media
An open server cabinet with several rack mounted servers, a pull out console tray with an LCD screen extended, and cable bundles running down the right hand side.
A cabinet with the door open and the console tray pulled out. That slide out screen and keyboard is a rack console, and it is how you work on a machine standing in front of it. Source: Wikimedia Commons, CC BY-SA 3.0.
Loading order, and the reasoning behind each choice

Bottom of the rack: storage and anything heavy. A fully populated storage node is genuinely heavy, and a tall cabinet with the weight at the top is a tipping hazard when you slide something out on rails. Same reason you put heavy books on the bottom shelf.

Middle: compute hosts. The bulk of the equipment, accessible without a ladder or crouching.

Top: network gear. The switch goes at the top so that every cable run from the servers below is short and travels upward, which is exactly where the name top of rack comes from.

Also consider the floor loading limit, which the colo provider publishes and which genuinely constrains dense storage builds.

Airflow discipline, and the cheapest mistake in this document

Blanking panels are not optional and they cost almost nothing

An empty U in the middle of your rack is an open hole between the cold aisle and the hot aisle.

Hot exhaust air takes the shortcut backwards through that hole into the cold aisle, where it gets inhaled by your servers. Your equipment now breathes its own hot exhaust, fans spin up, power draw rises, and component life shortens. All of it caused by a gap.

A blanking panel is a flat plate that fills an unused U and blocks that shortcut. They cost a rounding error and they are the highest return per dinar item in this entire document. Fill every unused U.

Rack elevation: the actual layout for our starter footprint

This is the artifact you hand to the install technician. It takes the Volume 0 equipment list and gives every item a specific U position.

U positionDeviceSizeWhy it is there
42Patch panel, copper1UTop, so all cable runs terminate above the gear
41Cable management bar1UDresses cables between panel and switch
40Top of rack switch, tor-sw-011UShort runs to everything below
39Out of band switch, oob-sw-011USeparate physical switch, isolated, section 8
38Cable management bar1U
36 to 37Edge firewall, fw-edge-012UNear the switch and the uplink handoff
35Blanking panel1UAirflow separation between zones
34cs-mgmt-011UCloudStack management server
33cs-mgmt-021UThe second one, for basic HA
32Rack console tray1USlide out screen and keyboard for on site work
28 to 31kvm-host-01 and kvm-host-022U eachCompute, in the accessible middle band
24 to 27kvm-host-03 and kvm-host-042U eachCompute, host 04 optional at launch
13 to 23Blanking panels11UReserved for growth, blanked today. Not left open.
9 to 12nfs-store-014UHeavy, so it sits low
1 to 8Blanking panels, growth reserve8USecond storage node lands here
Vertical, sidePDU A and PDU B0UMounted vertically in the side channels, section 3

Notice that the growth space is explicitly blanked rather than left open. Reserved does not mean empty. That is the difference between a plan and a hope.

Labeling, and the naming convention we commit to

Loading media
A cluttered equipment area with a black power strip covered in power plugs, each labelled with hand written masking tape, and unmanaged cables running across the floor to a UPS.
This is the failure mode. Hand written masking tape labels, unmanaged cables on the floor, and no way to trace anything without unplugging it. Every one of us has inherited something like this. Source: Wikimedia Commons, CC BY-SA 3.0.
Look at that photo, because it is the enemy

Masking tape and a marker at install time feels fine. Eighteen months later the adhesive has failed, the ink has faded, half the labels have fallen off, and the only way to identify a cable is to unplug it and see what breaks. In production. At night.

The standard we commit to:

Both ends of every cable get a printed label. Both. A cable labelled at one end only is a cable you will trace by hand.

Printed, not handwritten. A label printer costs less than one hour of remote hands time.

The label says where the other end goes, not what the cable is. kvm-host-01:eth0 on the switch end and tor-sw-01:Gi1/0/5 on the host end.

Device names match CloudStack host names exactly. If the machine is kvm-host-01 in CloudStack, that is the name on the physical label, in the DNS record, in the monitoring system, and in the wiki. One name everywhere. Two names for one machine is how a 3 a.m. incident becomes an hour of confusion about which box someone means.


3. Power, the thing that actually kills uptime

The analogy: the electrical panel in your house

You know intuitively that you cannot run the kettle, the oven, and a space heater on the same circuit, because the breaker trips. You know that a breaker tripping cuts everything on that circuit, not just the thing that caused it.

A rack is that, with a much bigger breaker and much worse consequences. And the mistake people make is thinking about power only as "will it turn on", when the real question is what happens when half of it goes away.

PDUs: the power strip that is not a power strip

A PDU, power distribution unit, takes the feed the colo gives you and splits it into outlets for your equipment. They come in three grades, and the difference genuinely matters.

TypeWhat it doesVerdict
BasicSplits power into outlets. That is all.Cheapest, and you are blind. You cannot answer "how much power am I actually using?"
MeteredReports actual current draw, per PDU or per outletBuy this. It pays for itself the first time you plan a capacity change.
SwitchedMetered, plus you can power cycle an individual outlet remotelyWorth it for the hard reboot you cannot do any other way. See section 8.
Why metered pays for itself, in one scenario

A customer asks whether you can host a workload that needs two more compute hosts. With a basic PDU your honest answer is "I think so", based on adding up numbers printed on the back of servers, which overstate real draw significantly.

With a metered PDU you read the actual current draw, compare it against your committed kW, and answer with a number. The second answer is a business capability, not a technical nicety, and it is the same argument as the utilization metric in Doc 2. Capacity you cannot measure is capacity you cannot sell.

A and B feeds, and the mistake I want you to never make

Serious equipment has two power supplies. The colo gives you two independent feeds, called A and B. The idea is that either one can carry the load alone, so losing a feed loses nothing.

Loading media
The rear of several rack mounted servers showing perforated power supply modules, network ports, a blue VGA connector, coloured cables plugged into each machine, and metal handles on the power supply units.
The back of real servers. The perforated grilles with handles are hot swappable power supplies, two per machine. The small ports beside them are network interfaces, including the dedicated management port from section 8. This is older gear, and the layout has not changed. Source: Wikimedia Commons, CC BY-SA 3.0.
The classic mistake, and it is genuinely common

Both power supplies plugged into the same PDU.

The server has two PSUs. The rack has two feeds. Everything looks redundant on the purchase order and on the diagram. And then feed A fails, and every server dies, because both of their power supplies were on feed A.

Dual power supplies protect you from a power supply failing, not from a feed failing, unless you actually split them. PSU 1 to PDU A, PSU 2 to PDU B, on every single device. It is like having two seatbelts bolted to the same weak point: it looks like redundancy and it protects against nothing.

This is worth checking physically, by hand, on install day, and it is on the verification list in section 10.

Loading media
Technical drawing of an IEC 60320 C13 socket and C14 plug showing their dimensions and pin arrangements.
C13 and C14, the connector pair on almost every server power cable. C14 is the male inlet on the equipment, C13 is the female connector on the cable. Source: Wikimedia Commons, CC BY 3.0.
Side definition: the connector alphabet soup

You will see these codes on every order form, and they are simpler than they look.

C13 and C14: the standard pair for most servers and switches, rated to 10 amps. The C14 is the male inlet on the device, and the C13 is the female end of the cable that plugs into it.

C19 and C20: the bigger version, rated to 16 amps, used by higher draw equipment like fully loaded storage nodes and some blade chassis.

Why you care: your PDU has a fixed number of each type. Order a PDU with only C13 outlets, buy a storage node that needs C20, and you have a box you cannot plug in. Count outlets by type, not just by quantity.

The power budget, done properly

Nameplate versus real draw, and why the difference matters

The nameplate rating is the number printed on the power supply. It is the absolute maximum that supply can deliver, and real servers draw substantially less, often around half of nameplate under normal load.

If you budget on nameplate you will conclude you need twice the power you actually do, and pay for it monthly. If you budget on idle draw you will trip a breaker the first time everything is busy at once.

The honest method: estimate real draw under expected load, add headroom, and then measure with a metered PDU and correct your numbers once the gear is running. This is why metered PDUs are in the budget.

Here is the shape of the budget sheet, using illustrative figures you replace with vendor data and then with measurements.

DeviceQtyEst. real draw eachSubtotalFeed
Edge firewall1100 W100 WA and B
Top of rack switch1150 W150 WA and B
Out of band switch130 W30 WA and B
CloudStack management servers2150 W300 WA and B
KVM compute hosts4450 W1800 WA and B
NFS storage node1500 W500 WA and B
Estimated totalabout 2.9 kW
With 30 percent headroomabout 3.8 kW
Per feed if one failseach feed must carry the full 2.9 kW alone

That last row is the one people skip, and it is the whole point of the exercise. If A and B each carry half the load in normal operation, then a feed failure means the survivor carries everything. So each feed must be sized for the full load, not half of it. A pair of feeds that can each carry 50 percent is not redundant, it is a synchronized failure waiting for a trigger.

Testing a feed failure without taking customers down

How to actually verify redundancy

Redundancy you have never tested is a belief, not a property of your system.

The safe test, run during a maintenance window while watching your monitoring:

  1. Confirm through the metered PDUs that both feeds are drawing current, which already catches the both PSUs on one feed mistake
  2. Confirm each device shows two healthy power supplies in its management interface
  3. Drop one feed, either by agreement with the colo or at your own PDU
  4. Watch for anything that powered off. That is your list of mistakes.
  5. Restore, wait, then repeat with the other feed

Step 4 is the entire value of the exercise. Almost every first test finds at least one device that was single fed, and finding it during a planned window is free while finding it during a real failure is an outage.

This feeds directly into capacity planning and the second rack trigger in executive metrics. Power, not rack units, is usually what runs out first. You will have empty U above a rack that has no watts left, and that is when you learn that a cabinet is sold by kilowatt, not by space.


4. Cabling and physical media

The analogy: plumbing

Copper is a garden hose. Cheap, flexible, easy to terminate, and it works fine over short distances. Push it too far and pressure drops.

Fiber is a light pipe. More expensive per end, immune to electrical interference, and it carries signal over distances copper cannot approach.

Here is the practical rule that covers ninety percent of decisions: inside the rack, use copper or DAC. Leaving the rack or the room, use fiber.

Copper: what the categories actually mean

Loading media
Close up of a green Ethernet patch cable with a transparent 8P8C modular connector, showing the eight individually coloured copper wires arranged inside the plug.
An Ethernet cable end. That connector is technically an 8P8C, universally called RJ45. Look through the clear plastic: eight wires in four twisted pairs, and the twisting is what makes the speed possible. Source: Wikimedia Commons, CC0.
Why the wires are twisted, since you can see it in the photo

Each pair is twisted around itself, and each pair uses a different twist rate.

Electrical noise hits both wires of a twisted pair almost identically. Since the signal is carried as the difference between the two wires, identical noise on both cancels out. The differing twist rates stop adjacent pairs from interfering with each other.

That is the entire trick behind gigabit and 10 gigabit over ordinary copper. It is also why you must not untwist more than about 13 mm at the connector, and why a badly terminated cable will pass a link light test and then fail under real load.

CategorySupportsPractical distanceUse in our rack
Cat5e1 Gbps100 mAcceptable for out of band and IPMI only
Cat61 Gbps, and 10 Gbps over shorter runs100 m at 1G, roughly 55 m at 10GThe sensible default for management and OOB
Cat6a10 Gbps100 m at 10GIf you want 10G over copper structured runs
Cat7 and aboveHigher, with caveatsVariesUsually not worth it. Go fiber or DAC instead.

The honest advice on copper categories: do not buy exotic copper to chase 10 gigabit. Inside a rack, DAC is cheaper, more reliable, and consumes less power than 10 gigabit over copper. Which brings us to the most useful cable in a hosting rack.

DAC: the cable that solved short 10 gigabit runs

Loading media
A black direct attach copper cable coiled in a loop, with a metal SFP form factor transceiver body permanently attached at each end, each having a pull tab release.
A DAC, direct attach copper. Twinaxial cable with the transceiver bodies permanently moulded onto both ends. You cannot separate them, and that is the point. Source: Wikimedia Commons, CC BY-SA 4.0.
Why DAC is the right default inside a rack

Look at the photo. Both ends are already transceivers. There is nothing to plug into anything, no optics to buy separately, no fiber end faces to keep clean.

Cheaper than two optical transceivers plus a fiber patch cable.

Lower power and lower latency, because there is no electrical to optical conversion happening at either end.

Nothing to contaminate. A dirty fiber end face causes intermittent errors that are genuinely painful to diagnose. A DAC has no end face.

The limitation: length. DAC is practical to roughly 3 to 7 metres, which is exactly the distance from any server in a cabinet to the switch at the top of that cabinet. Beyond the rack, you need optics.

Fiber and transceivers

Loading media
Side view of a small form factor pluggable transceiver module showing its metal casing, the electrical contact edge that inserts into the switch, and the optical port at the front.
An SFP transceiver. The contact edge on one side slides into the switch port, and the optical connector plugs into the front. This is the module that converts electricity into light. Source: Wikimedia Commons, CC BY-SA 4.0.
Side definition: the transceiver form factors

A transceiver is a removable module that converts your switch's electrical signals into light and back. The form factor names describe the physical shape and the speed.

SFP: 1 Gbps.

SFP+: 10 Gbps, same physical size as SFP. This is what our rack will use.

SFP28: 25 Gbps, still the same physical size.

QSFP+ and QSFP28: 40 and 100 Gbps, physically larger, four channels in one module.

Why the same size across speeds matters commercially: a switch with SFP+ cages can accept 1G or 10G modules in the same slot. You buy the switch once and change the optics as you grow, which is a genuinely useful property when budgeting.

Loading media
Labelled illustration of seven optical fibre connector types including FC, LC, FDDI, SC, SC-DUPLEX, ST, and MT-ARRAY.
The fibre connector zoo. You only need to recognize two of these in a modern rack: LC and SC. Source: Wikimedia Commons, public domain.
Loading media
Close up photograph of an LC optical fibre connector showing its small latching body and the precision ferrule that carries the fibre core.
An LC connector up close. Small, latching, and almost always used in duplex pairs because one fibre transmits and the other receives. This is what plugs into that SFP module. Source: Wikimedia Commons, CC BY-SA 3.0.
ConnectorRecognize it byWhere you meet it
LCSmall, with a plastic latch you squeezeAlmost everything modern. SFP and SFP+ modules use LC.
SCLarger, square, push and clickOlder equipment, and often the colo handoff at the patch panel
STRound with a bayonet twistLegacy. You may inherit it.
Single mode versus multimode, and the mistake that wastes an afternoon

Multimode fibre has a wider core, uses cheaper optics, and reaches a few hundred metres. Jackets are traditionally orange or aqua.

Single mode fibre has a very narrow core, uses more expensive optics, and reaches kilometres. Jackets are traditionally yellow.

The transceiver and the fibre must match. A single mode optic on multimode fibre may produce a link that appears to work and then throws errors under load, which is a genuinely miserable thing to chase. Match them deliberately, and ask the colo which type their handoff is on before you order optics, because guessing costs you a shipping lead time.

Patch panels, and when they earn their rack unit

Loading media
The inside of a rack showing three numbered patch panels at the top with short white patch cables looping down through horizontal cable management bars into two rackmount switches below, with a router at the bottom.
This is what good looks like. Numbered patch panels on top, switches below, and every cable routed through the horizontal management bars with a consistent service loop. Source: Wikimedia Commons, CC BY-SA 4.0.

Spend a moment on that photograph, because it is the target you are aiming for. Notice that every cable is the same length, they all follow the same path, and the loops are uniform. That is not vanity. It means you can trace any single cable with your eyes, and you can remove one without disturbing its neighbours.

Side definition: patch panel, and the wall socket analogy

A patch panel is a row of ports where permanent cable runs terminate. You then use short patch cables to connect a panel port to a switch port.

Why bother, instead of running the long cable straight into the switch? Because of the wall socket principle. You do not run electrical wire from the street directly into your lamp. You terminate it at a socket, then use a short flexible cord.

The permanent run gets terminated once and never touched again. All the plugging and unplugging happens on cheap, replaceable patch cables. Replacing a switch becomes moving patch cables rather than re terminating structured cabling.

Loading media
The rear of a tall stack of patch panels where dozens of grey cables are dressed in perfectly parallel horizontal bundles, secured with cable ties, converging symmetrically toward the centre of each panel.
The back of a structured cabling installation done properly. Someone spent real time on this, and every hour of it will be repaid the first time something has to be traced. Source: Wikimedia Commons, CC BY-SA 3.0.
Bend radius, the specification nobody reads

Every cable has a minimum bend radius, and violating it is how you create faults that appear months later.

For copper, a sharp bend distorts the twisted pair geometry and degrades performance at high frequencies. The cable still links, and it starts producing errors under load.

For fibre, bending too tightly causes light to escape the core. This is called macrobending loss, and it produces exactly the kind of intermittent, load dependent fault that takes days to find.

The practical rule: never pull a cable tight around a corner, never over tighten a cable tie to the point of deforming the jacket, and always leave a service loop, meaning a little slack, at each end. Look at the service loops in the two photos above. They are there on purpose.

Spare capacity, and the boring discipline that saves a growth day

ItemHow much spare to holdWhy
Switch portsAt least 25 percent free after installAdding a host should never require a switch purchase
Patch cables, each lengthA few of eachYou will damage one, and shipping takes days
TransceiversOne spare per type in useOptics do fail, and matching the exact type matters
DAC cablesOne or two spareCheap insurance
Blanking panelsMore than you thinkEvery device you remove leaves a hole that must be filled

The reasoning is simple: the marginal cost of a spare cable is trivial, and the cost of a growth request waiting three days on a shipment is a customer noticing. Spares are bought with the original order, not later.


5. Switching fundamentals for a hosting rack

Back to the receptionist, now with real behaviour

In Doc 1 a switch was a receptionist who memorizes which desk each person sits at. Here is what she actually does, in four steps, because these four steps explain almost every switching problem you will ever debug.

BehaviourWhat happensThe consequence you will see
LearningA frame arrives on a port, and the switch records "this source MAC lives on this port"The switch builds its table from traffic it observes, with no configuration
ForwardingDestination MAC is in the table, so the frame goes out that one port onlyThe normal, efficient case
FloodingDestination MAC is unknown, so the frame goes out every port except the one it arrived onTraffic appearing where you did not expect it. Normal, briefly, and a problem if constant.
AgingEntries expire after a few minutes of silenceWhich is why a quiet machine sometimes gets flooded to again

Here is the trick that makes flooding stop being mysterious: a switch only learns from traffic it sees. A machine that never sends anything is a machine the switch has never heard of, so every frame addressed to it must be flooded. This is also exactly why the ARP shout from Doc 1 works: the shout is the traffic that teaches the switch where the shouter lives.

Layer 2 switch versus layer 3 switch

Loading media
Front view of a 28 port rackmount gigabit Ethernet switch with two rows of RJ45 ports, status LEDs above each port, two SFP cages on the right, and rack mounting ears on each side.
A rackmount switch front panel. Count the RJ45 ports, then look at the right hand end: those two larger cages are SFP slots for the uplinks. Note the mounting ears, which is how it attaches to the rack rails from section 2. Source: Wikimedia Commons, CC BY-SA 3.0.
Layer 2 switchLayer 3 switch
Decides usingMAC addressesMAC addresses and IP addresses
Can route between VLANsNoYes
Has SVIsNoYes
CostLowerHigher
Side definition: SVI, and why it is the concept that unlocks this section

An SVI, switched virtual interface, is an IP address that belongs to the switch itself, inside a specific VLAN. It is a virtual port that lives in a VLAN rather than on a physical cable.

Why this matters enormously: an SVI is how a switch becomes the default gateway for a VLAN. Give the switch an SVI of 10.50.0.1 in the management VLAN, and every host in that VLAN can use 10.50.0.1 as its gateway, and the switch will route their traffic to other VLANs.

Without an SVI, a VLAN is an island. Hosts inside it can reach each other and nothing else. This is exactly the mechanism CloudStack expects for the management network gateway.

Top of rack design

Loading media
A row of populated server racks in a datacenter with numerous rack mounted servers, network switches, and bundles of cabling running through vertical cable management channels.
Populated racks in production. Every server in a cabinet runs a short cable up to the switch in that same cabinet, which is the entire idea behind top of rack. Source: Wikimedia Commons, CC BY-SA 3.0.

Top of rack means each cabinet has its own switch, and every server in that cabinet home runs to it with a short cable. Only the switch uplinks leave the cabinet.

PropertyWhy it wins
Short cablesCheap, DAC capable, easy to manage, easy to trace
Few cables leave the cabinetOnly uplinks cross between cabinets, so inter cabinet cabling stays sane
Fault domain matches the cabinetA switch failure affects that cabinet, which is a boundary you can reason about
Grows one cabinet at a timeAdding a cabinet means adding a switch, not re cabling the room
Our starter position, stated as the risk it is

At launch we have one top of rack switch. That switch is a single point of failure for the entire rack, and no amount of server redundancy behind it changes that. If it dies, everything in the cabinet is unreachable.

We write that down, monitor the switch properly, keep a configured spare if budget allows, and plan the redundant pair as an explicit upgrade. What we do not do is describe a single switch as a best practice. The honesty rule from Volume 0 applies: a documented risk is engineering, an undocumented one is negligence.

CloudStack's reference topology, and the collapse we make at our scale

CloudStack documents a specific reference shape, and it is worth stating precisely because Doc 5 depends on it.

CloudStack layerWhat the reference design expectsWhat we do in one rack
Pod level access switchA layer 2 switch that trunks all relevant VLANs to every hostOur ToR switch does this
Zone level layer 3 switchRoutes between networks and acts as the management network gatewayOur ToR switch also does this, via an SVI

At our scale these two roles collapse into one device, and that is fine, provided you understand that you have collapsed them. A layer 3 capable ToR switch serves as both the pod access switch and the zone gateway. The reason to know this is that when you read CloudStack documentation describing two separate switches, you should be able to map it onto your one box rather than assuming you are missing hardware.

The switch buying checklist

FeatureWhy it is non negotiable
802.1Q VLAN taggingSection 6. Without it there is no multi tenancy at all.
Trunk ports carrying many VLANsA hard CloudStack requirement, covered in section 6
SVIs and inter VLAN routingSo the switch can be the management gateway
LACP, 802.3adSection 7, for bonded host uplinks
MLAG or stackingNeeded later for a redundant pair that a single host can bond across
Jumbo frame support, 9000 MTUSection 7, for the storage VLAN
DHCP snooping and dynamic ARP inspectionThe ARP spoofing defence promised in Doc 1, enforced in Doc 4
RA guardThe IPv6 equivalent, promised in Doc 2
Port security and storm controlLimits the damage one misbehaving tenant can cause
sFlow or NetFlowTraffic visibility. You cannot investigate what you never recorded.
Dedicated management portSection 8. Managing a switch through the network you are reconfiguring is how you lock yourself out.
Serial console portThe last resort when everything else is unreachable

Oversubscription, in plain arithmetic

The restaurant analogy, and then the numbers

A restaurant has a hundred seats and a kitchen that can cook twenty meals at once. That is five to one oversubscription, and it works perfectly well because not everybody orders simultaneously.

A switch is the same. Suppose eight servers each connect at 10 Gbps, which is 80 Gbps of possible traffic into the switch. The switch has two 10 Gbps uplinks, so 20 Gbps out.

That is 4 to 1 oversubscription, and it is completely reasonable, because those servers are not all sending at full rate to the outside world at the same moment.

What actually matters is which traffic crosses the uplink. Server to server traffic inside the cabinet, including the storage traffic between hypervisors and the NFS node, never touches the uplink at all. So the ratio you care about is not total port capacity, it is the traffic that genuinely has to leave. Keep storage local to the cabinet and your oversubscription ratio stops mattering nearly as much.


6. VLANs: from concept to configured

The analogy: coloured stickers on the mail

You have one building, one corridor, and one mail trolley. But you have five departments, and department mail must never be delivered to the wrong department.

So you put a coloured sticker on every envelope. The trolley carries all colours down the same corridor. At each department door, only that colour is handed over.

A VLAN is the colour. The sticker is the 802.1Q tag. The corridor is the physical cable. And the whole point is that many separate networks share the same wire without ever mixing.
Side definitions: the three words that confuse everyone

Access port: a port belonging to exactly one VLAN. Frames arrive and leave untagged, because the device plugged in does not know VLANs exist. The department door.

Trunk port: a port carrying many VLANs, where frames are tagged so the other end can tell them apart. The corridor with the trolley.

Native VLAN: the one VLAN on a trunk whose frames travel untagged. The envelope with no sticker, where both ends have to already agree which department it belongs to.

The native VLAN, and why mismatches produce baffling failures

At first I found the native VLAN pointless, and here is the trick that made it click, along with the reason it is dangerous.

Untagged frames on a trunk have to belong to something. The native VLAN is that something. It exists so a trunk can also carry traffic from a device that does not tag.

Now the danger. Suppose the switch thinks the native VLAN is 1, and the server thinks it is 10. The server sends an untagged frame meaning "this is VLAN 10 traffic." The switch receives it untagged and concludes "this is VLAN 1 traffic." The frame is now silently in the wrong network.

Why this is a security issue and not just a bug

A native VLAN mismatch means traffic crosses a segmentation boundary without anything logging it or blocking it. There is no error. Both sides believe they are configured correctly. The frame simply arrives in a network it was never supposed to reach.

In a hosting environment where VLANs are how we keep tenants apart, that is a tenant isolation failure produced by a one line configuration disagreement.

The discipline: set the native VLAN explicitly on both ends of every trunk, use the same value everywhere, and never leave it at the vendor default. CloudStack reference designs commonly put hypervisor management on an untagged VLAN, which makes getting this right on the host trunks a launch requirement rather than a nicety.

The hard CloudStack requirement people miss

Every relevant VLAN must be trunked to every compute host. Including the public VLAN.

This is the single most common cause of a CloudStack deployment that installs cleanly and then behaves strangely, so read it twice.

The access switch must trunk all relevant VLANs to every compute host. Not just management and guest. The public VLAN too, on every host.

Why, when an external firewall handles guest gateways? Because CloudStack system VMs run on the hypervisors, and two of them need the public network directly:

The Secondary Storage VM needs public network access to fetch templates and ISOs from the internet.

The Console Proxy VM needs public network access so customers can reach their VM consoles from a browser.

Those system VMs can be started on any host in the cluster. So if even one host is missing the public VLAN on its trunk, everything works fine until a system VM happens to start on that host. Then console access or template downloads break, intermittently, in a way that looks like a CloudStack bug and is actually a switch port configuration.

Trunk every relevant VLAN to every host, identically. Uniformity is the requirement.

The VLAN numbering scheme we commit to

Numbers with meaning, and deliberate gaps for growth.

VLAN IDNameSubnetGatewayPurpose
10oob10.50.10.0/24OOB switch onlyBMC and IPMI. No route anywhere. Section 8.
20mgmt10.50.0.0/2410.50.0.1 on the ToR SVIHypervisor management, CloudStack management servers, system VM management
30storage10.50.1.0/24None, deliberatelyNFS traffic. Jumbo frames. Section 7.
40publicOur public rangeEdge firewallCloudStack public traffic. Trunked to every host.
50dmz10.50.5.0/24Edge firewallInternet facing platform services
100 to 999guest-pool-1Allocated per networkCloudStack Virtual RouterGuest VLAN pool. CloudStack allocates from this range.
1000 to 1999ReservedSecond guest pool, or a second zone
1Unused, on purposeVendor default, so we never use it for anything real
Two design choices in that table worth explaining

The storage VLAN has no gateway. This is deliberate, not an omission. Storage traffic never needs to leave its own VLAN, so giving it a gateway would only create a path for it to be reached from elsewhere. No gateway is a security control implemented by absence.

VLAN 1 is left unused. It is the default on essentially every switch, which means it is where mistakes land. An unconfigured port typically ends up in VLAN 1. If VLAN 1 carries nothing, then a mistake results in a port that reaches nothing, which is the failure mode you want.

Guest VLAN pool sizing, and the ceiling

The 4094 ceiling, and what it actually caps

A VLAN tag is 12 bits, which gives 4096 values, with two reserved. So you have roughly 4094 usable VLAN IDs, in total, across the whole layer 2 domain.

In CloudStack's isolated network model, each guest network consumes one VLAN. So the size of your guest VLAN pool is a direct cap on how many guest networks can exist simultaneously in that zone.

Our pool of 100 to 999 gives 900 guest networks. That is comfortably beyond the starter business plan and it is not unlimited. When you approach the ceiling, the answer is not a bigger VLAN range, because there is not one. The answer is VXLAN, which uses a 24 bit identifier and gives roughly 16 million segments.

VXLAN is the correct scale answer and it is not free: it adds encapsulation overhead that eats into your MTU, which is the subject of the next section, and it adds operational complexity. We start with VLANs, document the ceiling, and treat crossing it as a planned project rather than an emergency.

Disable VTP and its equivalents

Some vendors offer protocols that automatically propagate VLAN configuration between switches. They sound helpful and they are a genuine hazard, because a mistake on one switch propagates to all of them at machine speed.

We configure VLANs explicitly on each device, from version controlled configuration, in keeping with the Terraform and GitOps model from Volume 0. Explicit and boring beats automatic and surprising.


7. Bonding, redundancy, and MTU

The analogy: two ropes instead of one

You need to hold a weight. One rope might snap. So you use two.

Now there are two genuinely different ways to use two ropes. You can use one and keep the other slack as a spare, or you can share the load across both. The first is simple and wastes half your capacity. The second doubles your capacity and requires that both ends of both ropes are tied correctly.

That is exactly the choice between active backup bonding and LACP, and there is no universally right answer. There is only the answer that matches the switch you actually bought.
Active backup, mode 1LACP, 802.3ad, mode 4
How it worksOne NIC active, the other idle until the active one failsBoth NICs active, traffic distributed across them
Usable bandwidthHalf of what you installedAll of it, in aggregate
Switch configuration neededNone. The switch does not need to know.Required on both ends, and they must agree
Can span two separate switchesYes, triviallyOnly if the switches support MLAG or stacking
Failure behaviourVery simple and predictableFast, and more moving parts
Right choice whenYou have one switch, or two independent onesYou have an MLAG pair or a stack
Our starter decision, and the trap in it

We launch with one top of rack switch. Therefore we use active backup.

Here is the trap that catches people. With one switch, LACP would give us double the bandwidth and zero additional resilience, because both links terminate on the same box. If that switch dies, LACP does not save you. So LACP on a single switch buys throughput while adding a configuration dependency, and at launch throughput is not our constraint.

When the redundant pair with MLAG arrives, then LACP becomes correct, because at that point the two links genuinely terminate on two independent devices and the bond survives losing one.

Match the bonding mode to the switch topology, not to the bigger number in the datasheet.
Side definition: MLAG, and why it is the enabling technology

MLAG, multi chassis link aggregation, lets two physically separate switches present themselves to a connected server as if they were one switch.

That is what makes the good version possible. The server runs a single LACP bond, one link to each switch, and it does not know or care that they are two boxes. Lose a switch entirely and the bond keeps running at half capacity with no reconfiguration and no failover event on the host.

Vendors call it MLAG, VPC, MC-LAG, and other names. Whatever it is called, it must be on the buying checklist, because without it a redundant switch pair cannot carry a single bonded host.

Naming discipline, and the CloudStack requirement hiding in it

Bond and bridge names must be identical on every host in a cluster

This is not a style preference. It is a functional CloudStack requirement, and violating it produces a failure that appears much later and looks unrelated.

Live migration moves a running VM from one host to another. To do that, CloudStack must attach the VM's network interfaces to the same named bridge on the destination host.

If kvm-host-01 calls its guest bridge cloudbr1 and kvm-host-02 calls it cloudbr2, then everything works perfectly right up until the first migration attempt, which fails. Or worse, it succeeds and attaches to the wrong network.

Every host in a cluster uses identical interface, bond, and bridge names. Write them into the build automation so a human never types them, and verify them in section 10.

Verifying bond state and naming on a KVM host
root@kvm-host-01:~#

Read that first output like a checklist. Bonding Mode confirms you got the mode you intended. MII Status: up on both slaves means both cables are genuinely live, which catches the case where you built a bond on a cable nobody plugged in. Speed on each slave catches a link that negotiated down to 1 Gbps because of a bad cable. And Link Failure Count climbing over time is a cable or optic that is failing intermittently, which is exactly the fault you want to catch before it becomes an outage.

MTU, and the three layers you must plan separately

MTU came up in Doc 1 as the doorway height. Here are the actual numbers, and the thing to understand is that MTU is not one setting, it is three separate decisions.

WhereValueReasoning
Management VLAN1500Standard. There is no benefit to changing it and there is risk.
Public and guest VLANs1500Must be standard. This traffic goes to the internet, where 1500 is the assumption.
Storage VLAN9000, jumbo framesBulk NFS transfers, entirely inside our layer 2, measurable efficiency gain
Physical underlay, if VXLAN later9000, or at minimum 1600Must exceed guest MTU plus encapsulation overhead
Why jumbo frames help on storage, with the arithmetic

Every frame carries protocol headers. Those headers are roughly the same size regardless of how much data the frame carries.

With a 1500 byte MTU, moving a gigabyte of data takes on the order of 700,000 frames. With a 9000 byte MTU, the same gigabyte takes about 117,000 frames.

That is six times fewer frames, which means six times fewer headers to process and six times fewer interrupts for the CPU to service. On a storage network moving VM disk images and snapshots, that is a genuine and measurable reduction in CPU overhead and latency.

Why not use 9000 everywhere then? Because it only helps for large sustained transfers, and it requires every device in the path to agree. Internet traffic cannot agree, because you do not control the internet. So jumbo frames belong exactly where we put them: a dedicated, closed VLAN carrying bulk transfers between devices we own.

The absolute requirement, and how a mismatch presents itself

Every single device in the layer 2 path must be configured for the same MTU. Both hypervisors, the storage node, the switch, and every VLAN interface involved. One device left at 1500 breaks it.

And here is why this is the cruellest fault in the document. An MTU mismatch does not look like a network problem.

Small transfers work perfectly. Pings succeed. SSH is fine. Everything you instinctively test passes.

Large transfers stall, or run absurdly slowly, or hang and time out.

So the symptom is "our storage is slow" or "template downloads hang", and the team spends two days investigating disks, NFS tuning, and CPU load, because nobody suspects the network when ping works.

I want you to remember one sentence: if small things work and big things do not, check MTU first. It will save you a day at least once in your career.

Proving MTU end to end, the only test that counts
root@kvm-host-01:~#
How to read that test, because the flags are the whole point

-M do means do not fragment. This is essential. Without it the kernel quietly chops the packet up and the test passes even when the path cannot carry the size, which tells you nothing.

-s 8972 is the payload size. Add 28 bytes of IP and ICMP header and you get exactly 9000, the MTU. So this packet is the largest one that can possibly fit.

The first ping succeeding proves the entire layer 2 path carries 9000 byte frames. That is the real test.

The second command failing at 8973 is not a problem, it is the confirmation. It proves you found the exact ceiling rather than getting lucky with a smaller size. A test that only ever passes has not told you where the limit is.

Run this between every pair of devices on the storage VLAN, before CloudStack is installed. It is in the section 10 checklist.

VXLAN MTU arithmetic, at preview level

When VXLAN enters the picture, the overlay wraps every guest frame in additional headers, costing roughly 50 bytes.

So if the guest expects a normal 1500 byte MTU, the physical underlay must carry at least 1550, and in practice you want margin, so people commonly use 1600 as a floor or simply run the underlay at 9000.

The failure mode if you skip this is that guest VMs cannot pass full sized packets, which presents as the same maddening "small things work, big things fail" symptom, this time reported by customers rather than found by you. Full requirements land in Doc 5.


8. Out of band management, the lifeline

The analogy: the spare key with your neighbour

You leave a spare key with a trusted neighbour. Not because you expect to lock yourself out, but because the day you do, the alternative is a locksmith and a broken door.

Out of band management is that spare key, and here is the thing I want to be blunt about: you will need it. Not might. Will. Every engineer who has configured a network at any depth has, at least once, applied a change that severed their own connection to the device they were changing. The competent ones had a way back in.

Side definition: BMC, IPMI, iDRAC, iLO, and why one is a tiny computer

A BMC, baseboard management controller, is a small independent computer built onto the server's motherboard. It has its own processor, its own memory, its own firmware, and critically its own network port.

It runs whenever the server has power, even when the server is switched off. Think of it as the night watchman who stays awake in a closed building.

IPMI is the standard protocol for talking to a BMC. Vendors then brand their implementations: iDRAC on Dell, iLO on HPE, XClarity on Lenovo, IPMI generically on Supermicro. Same concept, different names and licence tiers.

What it lets you do, and this is the part worth internalizing: power the server on and off, watch the console as if you were standing in front of it including the BIOS and the boot messages, mount an ISO over the network to reinstall the operating system, read fan speeds, temperatures, and power supply health, and read the hardware event log to find out what failed.

All of that with the operating system completely dead. That is why it is called out of band: it is management that does not travel over the network the server itself uses.

Go back and look at the server rear panel photograph in section 3. Among those ports is a dedicated management port, physically separate from the data NICs. That single port is the difference between fixing a broken host from your desk and driving to the datacenter.

What OOB covers, and the small switch that serves it

DeviceIts OOB interfaceWhere it plugs
KVM compute hostsDedicated BMC port, IPMIoob-sw-01, VLAN 10
CloudStack management serversDedicated BMC portoob-sw-01, VLAN 10
Storage nodeDedicated BMC portoob-sw-01, VLAN 10
Top of rack switchManagement port, plus serial consoleoob-sw-01, VLAN 10
Edge firewallManagement port, plus serial consoleoob-sw-01, VLAN 10
The OOB switch itselfSerial console, and colo provided OOB if availableThe honest recursion problem, below
The recursion problem, stated honestly

If your OOB access runs through oob-sw-01, then what recovers oob-sw-01?

There is no clever answer. The chain has to terminate somewhere, and it terminates at something physical.

What we actually do: keep the OOB switch configuration deliberately trivial, essentially one flat VLAN and nothing clever, so there is almost nothing to break. Keep its configuration in version control so it can be restored quickly. Buy a serial console option from the colo if they offer one, which puts the last resort outside our own equipment entirely. And accept that in the worst case, the recovery path is remote hands with a laptop and a console cable, which is precisely why we asked about their response commitment in section 1.

Do not pretend this is solved. Know where the chain ends and how long that step takes.

Isolation, which is the part people get wrong

BMC firmware is a known weak point, and it holds total control

Be very clear about the security picture here, because it is uncomfortable.

A BMC can power cycle the machine, mount arbitrary media, and watch the console. Anyone who controls a BMC controls the server completely and beneath the operating system, where none of your host security tooling can see them.

And BMC firmware has a genuinely poor security history. Default credentials, patches that arrive late or never, and vulnerabilities that keep appearing. Attackers know this, and BMCs are a well documented lateral movement and persistence target precisely because they survive an operating system reinstall.

Therefore the isolation rules are absolute, not preferences:

Private addressing only, on VLAN 10, which appears in no public routing.

No route to the internet. None. Not for firmware updates, not for convenience. Updates arrive through a jump host.

No route to any tenant or guest network. Ever.

Reachable exclusively through a jump host or a VPN, with strong authentication and full session logging.

Default credentials changed on every device before it is racked, verified as a checklist item.

Never on a VLAN shared with anything else. A dedicated physical switch is cheap and the separation is worth more than the switch costs.

Here is the framing that makes this land for me: a compromised BMC is worse than a compromised operating system, because a reinstall does not remove it. It is the most privileged and least monitored surface in the rack. Treat it that way, and it becomes one of the controls we can actually demonstrate during the reviews described in security and trust posture.

The access runbook

QuestionOur answer
Who may reach OOB?A named, short list of infrastructure engineers. Not a role. Names.
How do they reach it?VPN, then jump host, then the OOB network. Two authenticated hops.
With what authentication?Multi factor at the VPN, keys at the jump host, individual credentials per device
What is logged?Every session, with who, when, which device, and what they did
How is access removed?Part of the leaver process, tested rather than assumed
What is the break glass path?Serial console, then colo remote hands, both documented with contact details and expected timings
The rehearsal that makes this real

A runbook nobody has practised is a document, not a capability.

Rehearse this deliberately: during a maintenance window, have an engineer deliberately break the network configuration on a non production host, then recover it using only OOB. Time it. Write down every step that was unclear, every credential that was missing, and every assumption that turned out to be wrong.

Do it before you have customers, because the alternative is discovering the gaps during a real incident with an audience.


9. The physical network zone map, finalized

Volume 0 named the zones. Doc 5 will consume them. This is the authoritative table in between, and from here onward this is the version everything else references.

ZoneVLANSubnetGatewayAddressingMTUWhat connects
WAN uplinkUntagged on the handoffAssigned by the carrierCarrier sideStatic, from the carrier1500Colo handoff into the firewall outside interface
Edge and DMZ5010.50.5.0/2410.50.5.1, firewallStatic1500Internet facing platform services
Management2010.50.0.0/2410.50.0.1, ToR SVIStatic for infra, CloudStack pool for system VMs1500Hypervisor management interfaces, both CloudStack management servers, system VM management addresses
Public40Our allocated public rangeFirewallCloudStack public IP pool1500Trunked to every host. Virtual Routers, SSVM, CPVM.
Guest100 to 999Per guest networkCloudStack Virtual RouterVirtual Router DHCP1500Customer VM traffic, one VLAN per isolated network
Storage3010.50.1.0/24None, deliberatelyStatic only9000Hypervisor storage interfaces and the NFS node
Out of band1010.50.10.0/24None reachable from productionStatic1500Every BMC, plus switch and firewall management ports
Two CloudStack requirements encoded in that table

Private and public networks require separate subnets. CloudStack advanced zones require distinct subnets for the private and public networks. This is not a recommendation you can economize on, and a deployment that shares them behaves incorrectly.

Storage needs its own subnet, and here is the actual reason. People assume the dedicated storage VLAN is purely about performance. Performance is a benefit. The requirement is that a hypervisor with two interfaces needs a way to decide which interface to send storage traffic out of, and it makes that decision from the routing table, which means the storage network must be a different subnet from the management network.

Put storage and management in the same subnet and the hypervisor will send storage traffic out whichever interface the routing table happens to prefer, which is usually the management one. Everything appears to work, your jumbo frames do nothing, and your storage traffic competes with management traffic. The separate subnet is what makes the interface separation real.

The buildable diagram

This replaces the executive sketch from Volume 0 with the version you could hand to an installer.

Preparing diagram

The three things to read off that diagram: every host trunk carries the identical VLAN list including 40, which is the CloudStack requirement from section 6. The OOB paths are dotted and terminate nowhere else, which is the isolation from section 8. And both power feeds enter the rack, which is section 3.


10. Install day and verification

The analogy: the pre flight checklist

A pilot with thousands of hours still reads the checklist aloud before every flight. Not because the knowledge is missing, but because under time pressure, humans skip steps, and the checklist is what catches that.

Install day is time pressured by definition. You have a booked window, possibly a drive or a flight, and a technician waiting. This is exactly the situation where a checklist outperforms competence.

Why verification before CloudStack is not optional

The CloudStack community's own troubleshooting guidance says something worth repeating: in the vast majority of cases where a deployment misbehaves, the problem turns out to be the switching layer configured incorrectly.

Think about what that means for how you spend install day. If you install CloudStack on an unverified network, then every subsequent problem has two candidate causes, and you will spend your time reading CloudStack logs hunting for a bug that is actually a trunk port.

So we prove the network works while the network is the only thing that exists. Every test below runs before CloudStack is installed. Then when CloudStack does misbehave, you can genuinely say the network is not the cause, because you hold the evidence.

Everything on this list is far easier at your desk than in a cold aisle.

  • Firmware updated on BIOS, BMC, NICs, and storage controllers, then recorded so every host matches
  • BMC configured with a static address on VLAN 10, and default credentials changed
  • BIOS settings applied identically on every host: virtualization extensions enabled, power profile set, boot order set
  • Interface naming confirmed, so you know which physical port is which logical name before you cable anything
  • Drives installed and the storage controller configured
  • Labels printed for every device and every cable, then packed with the gear
  • Configuration files staged in the repository for the switch, the firewall, and the host build
The one that always gets skipped
Change the BMC default credentials before the machine is racked.

Once it is racked and you are behind schedule, this becomes a task for later, and later becomes never. A BMC sitting on your management network with vendor default credentials is the single most valuable thing an attacker can find, for exactly the reasons in section 8.

The as built document

What we record, and where it lives

The rack does not match the plan. It never does. Something was out of stock, a port was faulty, a cable was too short.

The as built document records what is physically there, and it gets written on install day while the details are in front of you, rather than reconstructed later from memory.

Contents: the final rack elevation with actual U positions, the cabling schedule with actual ports and labels, serial numbers and asset tags per device, firmware versions as installed, the final VLAN and subnet table, IP assignments including every BMC, the power feed assignment per device, and photographs of the front and rear of the rack.

Where it lives: in the repository, version controlled, beside the switch and firewall configurations. Not on someone's laptop, not in a shared drive nobody can find, and not in the head of whoever did the install. That is the zero tribal knowledge commitment from Volume 0, applied to hardware.

Take the photographs. Genuinely. A clear photo of the front and the rear of the rack, taken on install day, answers questions for years afterwards and costs thirty seconds. Every time somebody asks which port host 03 is in and you answer from your desk instead of calling remote hands, that photo has paid for itself again.


Labs and deliverables

DeliverableFormat
Rack elevationU by U layout with device names, weights, and power draw. Section 2 gives you the template.
Power budget sheetPer device draw, totals, headroom, and A and B feed assignment, with the single feed survival row
Cabling scheduleSource device and port, destination device and port, media type, length, and the exact label text for both ends
VLAN and subnet tableThe authoritative version from section 9, consumed by Doc 4 and Doc 5
Switch configuration outlineTrunk ports, access ports, native VLAN, LACP groups, SVIs, and jumbo frame settings
Trunk verification testTagged interface ping plus the tcpdump capture proving the tag, per VLAN per host
MTU verification recordThe do not fragment ping results between every pair of devices on the storage VLAN
Hardware requirements listSwitch, firewall, transceiver, and cable specifications for procurement, built from the section 5 checklist
OOB access runbookThe section 8 table filled in with real names, real paths, and a rehearsal date

Honest tradeoffs we are documenting, not hiding

ChoiceStarter positionScale positionWhat triggers the change
ToR switchingSingle switch, documented as a single point of failureRedundant pair with MLAG and LACP to hostsFirst paying customer with an availability commitment
Edge firewallSingle applianceHigh availability pairSame trigger as the switch pair
Host uplinksTwo NICs bonded active backup, management and guest separatedFour NICs in two bonds, or 25G with overlayUplink utilization, or the arrival of an MLAG pair
Tenant isolation transportVLAN per guest networkVXLAN with EVPNApproaching the guest VLAN pool ceiling
Storage transportDedicated storage VLAN, NFS, jumbo frames, one nodeHigher speed links, possibly Ceph with its own considerationsStorage IOPS or capacity metrics, not ambition
PDUsMetered, at minimumSwitched, for remote power cyclingThe first time you need a hard reboot and cannot get one

We write the starter position down as a known risk, not as a pretend best practice. That is the same honesty rule Volume 0 used for the single firewall and single ToR, and it is the reason the tradeoff table has a trigger column. A risk with a documented trigger is a plan. A risk without one is just an excuse that ages badly.


Success criteria

You are done with this document when you can:

  • Recognize and name every component in the photographs above, and explain what each one is for
  • Hand a colo provider a rack elevation, a power request, and a cross connect request without hesitating
  • Explain why every VLAN must be trunked to every hypervisor, including the public VLAN, and name the two system VMs that make it necessary
  • Choose between LACP and active backup and defend the choice using the switch you actually bought
  • State the MTU on each VLAN and prove it end to end with a do not fragment test
  • Describe, step by step, how you regain control of a host whose network configuration you just broke
  • Explain why the storage VLAN needs its own subnet, and why it has no gateway
The test I would actually give you

Forget the list for a second. Here is the real one.

Someone hands you a photograph of the back of a rack. Can you name what you are looking at? The power supplies, the management port, the data NICs, the DAC cables, the transceivers, the patch panel, the cable management bars, and whether the person who built it split the power feeds.

If you can do that, you have the thing that separates someone who has worked with this hardware from someone who has read about it. That was the entire point of putting the photographs in.

This document is where the handbook stops being theory. Everything after it assumes a rack that a professional would be comfortable owning, and a set of decisions you can defend in a room with a vendor, an auditor, or a customer.


Next: Segmentation, firewalls & edge defense. The wires exist and the VLANs are defined. Now we decide who may talk to whom, and enforce it.