Cledion
CledionStratum manual
Stratum Manual

Operate your data center from one screen.

A practical guide to running Cledion Stratum in your facility: read the Operations Overview, drill into a rack's power and cooling, tune your alerts, build the screens and Wall your team lives on, and pull the reports management asks for.

Book a walkthrough For operators & NOC teams
On this page
01Getting oriented

What Stratum is

Cledion Stratum is the monitoring platform for your data center. It polls the gear on your floor, switches, routers, firewalls, servers, PDUs, UPSes, and environmental sensors, over SNMP, stores the telemetry on-site, and turns it into live status, history, alerts, and per-rack power and cooling analytics for your operations and NOC teams.

It is a DCIM built for ChilliRack direct-rack-cooling deployments, so power and cooling are first-class, not an afterthought bolted onto a network tool.

The platform is built around one idea: a facility produces thousands of raw readings, but an operator only needs the few that change a decision. Stratum collapses 687 raw sensors into the roughly 5 decisions that actually need a human on a given shift.

Who it's for

Data-center operators and NOC teams who need real visibility into power, cooling, and network across the floor, without shipping telemetry to someone else's cloud.

Stratum is read-only. It observes, charts, alerts, reports, and advises. It never sits in the control path of your power, cooling, or network, so turning it on cannot change how the facility runs.

02Getting oriented

How it works

You do not need the internals to use Stratum, but a few concepts explain what you are looking at.

On-prem · telemetry never leaves the facility
On-prem core
Stratum + time-series store
SNMP polling · syslog · service checks
Switches
Servers
PDUs
UPSes
Sensors
Stratum runs on-prem, polls the racks around it, and keeps every reading inside your facility.
  • Agentless polling: Stratum reads your switches, PDUs, UPSes, iDRACs, and sensors over SNMP GET. Nothing is installed on the monitored gear.
  • One probe per rack or row: a lightweight probe sits next to the hardware and pushes readings back to the core, so segmented and remote rows are covered without inbound firewall rules.
  • On-prem and read-only: the core runs inside your facility and telemetry never leaves it. Stratum is advisory, never in the control path of power, cooling, or routing.
  • Optional advisory AI: an on-box advisor can summarize conditions and flag what deserves attention first. It is off by default, holds no credentials, and makes no outbound calls. Turn it off and Stratum is a zero-AI system.
03Getting oriented

The interface at a glance

The left sidebar groups everything Stratum does into a few areas. This manual is organized the same way, so where you are in these pages maps to where you are in the app.

AreaPagesWhat it answers
OverviewOperations OverviewIs the facility healthy right now?
RacksPer-rack viewHow is each rack drawing power and shedding heat?
Power & coolingPower & Energy · Thermal & CoolingAre we efficient and inside the envelope?
NetworkNetwork · SyslogWhat is the network doing, and what just changed?
InventoryDevices · Sensors · ProbesWhat are we watching, and how?
Control roomScreens · Wall · AlertsWhat needs attention, and where?
OperationsReports · SettingsRecords and configuration
04The operations picture

Operations Overview

The Operations Overview is the fast-scan view for a shift: facility-wide health, live power draw, cooling status, active alerts, and the KPIs that matter across the floor on one screen. It is built for a glance during normal operations and for quick triage during an incident.

  • Facility totals: IT load, total power, and PUE, trended so you see drift, not just the instant value.
  • Cooling at a glance: how many racks sit inside their thermal envelope and which are creeping toward the edge.
  • Active alerts: the open items a human still needs to handle, most urgent first.

Every card drills down. Click a rack, a PDU, or a sensor and you land on the detailed view for it, so the overview is also the front door to everything else.

05The operations picture

The per-rack view

The per-rack view is where a single rack tells its whole story: power in, heat out, and how hard the cooling is working to keep up.

Rack PUE

Each rack reports its own PUE, computed from the metered feeds and the ChilliRack fan: PUE = (PDU1 + PDU2 + fan) / (PDU1 + PDU2). A rack drawing 10 kW of IT and 0.5 kW of fan power runs a rack PUE of 1.05. Watch it over time: a rising rack PUE at steady IT load usually means the cooling is working harder than it should.

Thermal delta (Delta T)

Delta T is the temperature rise across the rack, return air minus supply air. A healthy, well-loaded rack holds a steady Delta T. A collapsing Delta T can mean bypass air or an over-provisioned fan, and a climbing one can mean the rack is out-running its cooling.

Cooling efficiency (CER)

The cooling-efficiency ratio is IT load divided by fan power: CER = IT load / fan power. It answers a simple question: how many watts of compute am I cooling per watt of fan? Higher is better. Read rack PUE, Delta T, and CER together and you can tell whether a hot rack is a load problem or a cooling problem.

06The operations picture

Power & energy

The Power & Energy page rolls the facility's electrical picture up from the individual PDU feeds.

  • PDU breakdown: draw per PDU and per phase, so you can spot an unbalanced feed or a rack approaching its breaker before it trips.
  • Energy and cost: kWh over the day, week, and month, converted to cost at your tariff so power shows up in the units management cares about.
  • PUE history: facility PUE trended over time, not just a single headline number, so efficiency work is something you can actually see.
Metered, not modeled

Power figures come from the metered PDU feeds Stratum polls, not from nameplate estimates. What you see is what the racks actually drew.

07The operations picture

Thermal & cooling

The Thermal & Cooling page shows inlet and return temperatures across the floor and how each rack sits against its operating envelope.

ASHRAE bands

Inlet temperatures are drawn against the ASHRAE thermal guidelines: the recommended band, roughly 18 to 27 °C (64 to 81 °F), and the wider allowable band for your equipment class. A rack inside the recommended band is comfortable, one in the allowable band is fine but worth watching, and one past it needs attention. The bands sit right on the charts, so you read temperatures in context rather than against a number you have to remember.

Because Stratum knows both the inlet temperature and the fan power behind it, the thermal view and the cooling-efficiency numbers on the rack view tell one consistent story. You can chase a hot inlet back to the rack that is causing it.

08The operations picture

Network & syslog

The Network page covers the connectivity side of the floor: interface throughput, link state, and device reachability, alongside the syslog stream where devices report what just happened.

Reviewing syslog next to the metrics gives you a second angle on an incident. A link flap or a firewall policy reload often explains a graph before any threshold trips.

A note on syslog

Syslog runs over UDP, which has no handshake, so a source address can be spoofed on a reachable network. Stratum binds sources and caps per-source floods to keep unregistered chatter off the page, but treat syslog as a signal, not proof of origin.

09Devices, sensors & alerts

Devices & sensors

Everything Stratum watches lives in the inventory: the devices on your floor and the sensors on each one. A device is a piece of gear, a switch, a PDU, a server, an environmental sensor. A sensor is one measured series on it, such as an interface's throughput or a rack's inlet temperature.

Add a device

DevicesAdd device
  1. 1Name it, then choose a type (switch, router, firewall, server, PDU, UPS, sensor), enter its IP address, and place it in a group and, optionally, a rack.
  2. 2For SNMP gear, set the credentials (v1/v2c community or a v3 user) and test connectivity before saving.
  3. 3The device appears immediately in the inventory, on the topology map, and on the status wall.

Add a sensor

Devicesexpand the deviceAdd sensor

A sensor is named device_id.sensor_name. Name it in snake_case by convention, inlet_temp_f, power_w, ping_gateway_ms, and set its unit. Follow the temperature suffix _f or _c and the °F/°C toggle converts the series for you.

Probes register themselves

You rarely add gear by hand. A remote probe auto-registers the devices it polls, so most of your inventory populates itself as you point probes at rows.

10Devices, sensors & alerts

Remote probes

Remote probes are how Stratum reaches across a floor without a flat network. A probe is a lightweight collector that sits next to a rack or row, polls the gear around it, and pushes the readings back to the core.

Stratum core · one per site
Core service · time-series store · GUI
outbound HTTPS · HMAC-signed · revocable keys
Probe · Rack A
SNMP GET only · never writes to a device
Probe · Rack B
SNMP GET only · never writes to a device
One core per site; a probe sits next to each rack or row and pushes signed readings back over outbound HTTPS.
  • One probe comfortably covers a rack or a row on a standard polling interval; add probes as the device count grows.
  • Probes push back over outbound HTTPS, so they reach subnets behind NAT with no inbound firewall rules.
  • Every push is signed and each probe key is scoped and revocable. A probe does SNMP GET only and never writes to a device.
  • On first ingest the core auto-registers the probe's devices, with no manual steps.
11Devices, sensors & alerts

Alerts & alert rules

Alert rules do the watching so nobody has to stare at a feed. You set a threshold on the readings that matter, and Stratum raises an alert when a value crosses it.

Control RoomAlerts
  1. 1Create a threshold rule on a sensor, for example inlet temperature above 27 °C, or a PDU feed above 90 percent of its rated draw.
  2. 2When the value crosses the limit, an alert is raised and surfaced on the Alerts page and the Wall.
  3. 3Acknowledge an alert to show the team it is being handled, and resolve it once the condition clears.
Known vs. new

Acknowledging and resolving is what separates known issues from genuinely new work, so the Alerts page stays a to-do list rather than an undifferentiated stream. This is the 687-readings-to-5-decisions idea in practice: rules filter the noise so the page only shows what a human still needs to do.

12Screens, walls & reports

Screens & the Wall

Stratum ships with sensible default views, but the control room is where you make it yours.

Screens

A Screen is a dashboard you compose from widgets: pick the charts, rack tiles, alert lists, and KPI cards you care about and arrange them for your role. A power engineer's screen and a network operator's screen can look nothing alike and both be right, and each operator can keep their own.

The Wall

The Wall is a full-screen, glanceable display for the shared NOC monitor: facility health, active alerts, and the racks under stress, sized to read from across the room. It is the same data as everything else, drawn for a wall instead of a desk.

One source, many walls

A single core can drive many walls and screens at once. The display layer is just a view onto the same core, so every screen in the room agrees on the numbers.

13Screens, walls & reports

Reports & exports

Reports turn monitoring history into something you can hand off. Stratum generates operational reports on a schedule and lets you export the underlying data for deeper review.

  • Daily and monthly reports: a scheduled PDF summarizing power, cooling, PUE, and the alerts that fired, ready to send to a teammate, a vendor, or management.
  • CSV export: pull the raw readings for any device or sensor over any window, for your own analysis or for an audit trail.
Written for handoff

Reports are built to be read by someone who was not on shift. The point is that the record explains itself without you in the room.

14Safety & the advisory AI

The advisory AI

Stratum includes an optional, on-box advisor: explainable analytics that summarize conditions and point you at what deserves attention first. It runs as a separate service, holds no credentials, makes no outbound calls, and anything it builds stays inside your facility. It is off by default, and with it off Stratum is a zero-AI system.

The model proposes, deterministic code disposes

Where Stratum uses AI, intelligence stays probabilistic and safety stays deterministic. Anything the advisor suggests is a proposal that deterministic rules approve or refuse, and nothing happens without a human. An agent can never acknowledge a critical alert or disable a rule on its own.

Agent proposes
a schema-validated proposal
Deterministic rules
approve or refuse
A human acts
nothing runs without you
Physical control plane
Power · cooling setpoints · routing. AI never touches this.
Intelligence stays probabilistic; safety stays deterministic, and the physical control plane is off-limits.

One line is permanent: AI never touches the physical control plane. There is no endpoint anywhere in Stratum that switches power, changes a cooling setpoint, or edits routing. The platform runs the monitoring; people and physical systems keep control of the facility.

Was this page helpful?

Ready to see it on your own floor?

Book a walkthrough and we will show you Stratum reading real power, cooling, and network telemetry, all of it running on-prem inside your facility.

Book a walkthrough