Boxing Day, 03:40
The power went off
The machine is fine. It is simply not running, and it will not be running until somebody walks into a closed building and presses a button.
Dead man’s switch monitoring · servers, NAS and PCs
Every few minutes, each machine you look after says “I am still here” to a hub outside its building. Red Kite watches for the message that does not arrive, and tells a person, in a sentence, within minutes.
Free while in beta · unlimited machines · you host it · no account
Nothing to open on your firewall. No VPN, no port forwarding, no static IP. A server in a cupboard behind a domestic router is as watchable as one in a data centre.
answering late past 5m silent past 15m
A machine that has just reported sits at the start of the scale and drifts clockwise as it goes quiet. A tick crossing a mark is a machine changing state: visibly, in the direction of trouble.
Ashford server is late, nothing for 12:34, but it is still inside its grace period, so nothing has been raised.
The gap
A monitor that lives on the same site, on the same power, behind the same broadband, cannot report its own absence. It goes quiet at exactly the moment you need it, and silence from a monitor reads like a quiet night.
Boxing Day, 03:40
The machine is fine. It is simply not running, and it will not be running until somebody walks into a closed building and presses a button.
Any Tuesday
Everything on site is healthy and nothing on site can tell anybody. The router dialling home is the router that stopped.
The August fortnight
An office closed for two weeks with a server quietly dead in a cupboard is the classic. It is discovered on the first morning back.
The one nobody plans for
No alerts is what a dead monitor and a good week both look like. That ambiguity is what quietly kills systems like this, so we answer it three ways.
A monitor that must reach in reports “all well” when its own network breaks. This one reports trouble when anything between the machine and the hub breaks, which is the honest answer, because if we cannot hear from it, we do not know that it is well.
The inversion
This single decision is why Red Kite works in places ordinary monitoring cannot go, and it is a decision we will not trade away later for convenience.
Ordinary monitoring polls: it sits somewhere and connects to each machine to ask how it is. That needs a route in, a port forwarded, a VPN, an agent listening on something. At eighty customer sites behind eighty domestic-grade routers, that is eighty negotiations, eighty firewall rules and eighty things that can be got wrong.
Red Kite inverts it. The machine makes one outbound HTTPS request every few minutes, the way it already reaches every website and update server it uses. Nothing is exposed, nothing is forwarded, and there is no new attack surface at the customer’s edge, because there is nothing new listening.
It also means we are honest about the limit this buys. We cannot tell “the machine is off” from “the broadband is down”, and we do not pretend to. We say what we know: we have not heard from this since 04:12.
The screens
Worst first, always. The sentence across the top is the product; the figures underneath support it. This is the Now screen, drawn with the hub’s own components and an example fleet of fourteen machines.
We have not heard from Ashford server since 04:12 this morning.
10 answering. 2 silent, 1 late but inside grace, 1 never heard from.
Every tick is a machine, placed by how long it has been quiet.
We have not heard from Ashford server since 04:12 this morning.
AcknowledgeWe have not heard from Selby PC since 04:23 this morning.
Example figures on an example fleet. The machine names, customers and times are made up; the layout, the components and the wording are the hub’s own.
Every colour has a shape
Silent is a diamond, late is a triangle, answering is a circle, and never-heard-from is a hollow square. Red and green are exactly the pair colour blindness destroys, and this screen may hang on a wall, so identity never rests on hue alone.
One badge per kind of trouble
“This server is unwell” is the beginning of a question, and a single worst-of badge asks somebody to open the machine to answer it. “Its disks and its backup are unwell, and its services are fine” is the beginning of a decision. So each part of a machine gets its own mark, with a count where there is more than one.
Grey is never a pass
A machine we have never heard from gets the same visual weight as one that has stopped, a different colour, the same insistence. It is not a faint tick waiting to go green. We cannot tell whether it is well, and the screen says so.
The second job
A server with a failed mirror leg, a filesystem quietly remounted read-only and four hundred corrected memory errors answers every heartbeat perfectly: right up until the morning it does not answer at all. These are the failures that present fine in the foreground and kill the machine in the background.
When ext4 hits an I/O error it remounts read-only to protect what is left. Everything already running carries on from memory. Nothing crashes. The site still serves.
Why nothing else sees it: every write silently fails until somebody reboots days later and finds that the last week does not exist. A machine in this state passes every heartbeat check ever written, including ours.
The quietest catastrophe there is. A RAID 1 with one disk gone serves every read perfectly and gives no sign at all from the foreground.
Why nothing else sees it: the machine is fine until the surviving disk goes, and then there is nothing. We read mdadm and ZFS state, rebuild progress, and the mismatch count, which on RAID 1 is silent corruption nothing else mentions.
A DIMM correcting single-bit errors is doing its job and telling nobody. We read the EDAC counters on Linux and WHEA events on Windows.
Why it matters: a module that has thrown one correctable error is far likelier than average to throw an uncorrectable one, and that is a machine that stops dead in the night with no explanation attached.
An LVM thin pool that reaches 100% does not return an error politely. It wedges every filesystem sitting on it, at once.
Why nothing else sees it: not a single disk is full. Every tool that measures free space says the machine is fine, right up to the moment everything stops.
A backup is judged on when it last succeeded, never on what the last run said. Failed systemd units are reported by name, “one unit failed” is not actionable; borgbackup.service is.
Why nothing else sees it: a backup failing quietly is only ever discovered on the day it is needed.
Windows disk events 129 and 153 are a controller resetting a device that stopped answering, and a retried read. On Linux, the SMART command-timeout and pending-sector counts say the same thing.
Why nothing else sees it: logged as warnings, surfaced by nothing. The machine is perfectly usable for months. It just occasionally stops, and nobody can say why.
A filesystem can run out of inodes while reporting hundreds of gigabytes available. It is full. Nothing can be written.
Why nothing else sees it: every tool that measures free space in bytes says it is fine, because in bytes it is.
A kernel security update installed six weeks ago on a machine nobody has restarted is not installed. We also derive unplanned restarts from uptime going down between two check-ins.
Why nothing else sees it: the update counter reports zero outstanding, which reads as fully patched and is not. Nothing anywhere records a reboot.
The principle behind the figures
A drive with twenty-four reallocated sectors that has had twenty-four for three years is a scar. The same twenty-four appearing this week is a drive coming apart. Same number. Opposite meanings. Only a comparison tells them apart.
Every counter rule compares the newest check-in against one from seven days ago, and says on screen which of the three cases it is looking at: rising, steady, or not yet knowable.
The same reasoning runs through the space rules. A disk at 82% is not interesting. A disk at 82% that was at 61% last week runs out a week on Thursday, and that is a date somebody can put in a diary and a job somebody can quote for.
We take the raw SMART value, never the vendor’s normalised one. Vendors disagree about what “100” means; only the raw count can be compared against last week’s.
The rule we will not bend
A monitoring system that renders “cannot tell” as “fine” is worse than having none at all, because it is believed. Wherever a reading is missing, stale or unreachable, we say so, in its own colour, with words beside it.
Nailsea server
41% free, and steady for a month.
Tring NAS
Nearly full: 3% free, and backups stop when it fills.
Melksham office PC
We cannot read this disk, so we cannot say whether it is well.
That third card is the whole rule. An empty ring resting at zero and a healthy zero are the same picture, and the difference between them is “we cannot see the disk” against “the disk is fine”. So the unknown state is drawn differently, and the words are written next to it.
The same applies everywhere else: a machine that has never checked in is grey with an explanation, not green. A machine whose last check-in is stale is amber or red on the strength of that staleness alone, because the figures inside a stale check-in describe the past. And a snapshot older than two hours is discarded rather than sent: vanishing is honest; freezing at last week’s values is the exact lie this product exists to catch.
The loop, precisely
It gathers its figures, every one a cheap read of /proc, df or the Windows equivalent, and posts them to the hub with its own token. No disk scans, no database queries, no process trees.
Machines have wrong clocks, across eighty of them there are always some. The agent’s own clock travels as a fact, so we can tell you a machine is 41 minutes behind, but it is never the basis of a decision.
It finds every check whose last check-in is older than its interval plus its grace period. One query on an indexed timestamp, which is simple enough to reason about at three in the morning, and which catches up correctly if the hub itself was off for an hour.
Not one message per sweep. An incident opens once, notifies once, and escalates on a schedule. Fifty texts about one dead server is how somebody ends up blocking the number.
On every channel they have chosen for that severity. Unacknowledged after a set time, it escalates, the next person, a louder channel, or both. Acknowledging stops the ladder, and needs no sign-in.
The commonest ending is the machine returning on its own, and a resolution message goes out on the same channels as the alarm. Leaving people to tidy up stale incidents by hand is how the list stops being believed.
Getting the message through
Alerts go by every channel a person chooses, per severity. Not one channel configured globally for everybody. The standard is “no excuse to miss all three”.
| Channel | Info | Warning | Critical |
|---|---|---|---|
| Push | Yes | Yes | Yes |
| No | Yes | Yes | |
| Text message | No | No | Yes |
| Quiet hours | 22:00–07:00, overridden for critical only | ||
Quiet hours have a per-severity override. Critical still wakes you at two in the morning; a disk at 82% waits until breakfast. An engineer who looks after four accounts is not woken for the other seventy-six.
Not sending is a failure, and it is recorded as one. Every attempt is stored with its outcome and the provider’s own reference. If a channel fails, the next is tried; if all of them fail, that itself becomes an incident. A monitoring system that cannot prove it sent the message has not sent it.
Email goes through a transactional provider with SPF, DKIM and DMARC set properly, not a mail relay with its own silent failure modes. The wrong dependency for the thing that tells you about failures is another thing that fails quietly.
The hardest problem here
Silence is ambiguous, and it is ambiguous in the worst possible direction. Three answers, all cheap, all used together.
Every morning, 07:00
“All 43 machines answered overnight. Nothing outstanding.” The value is not the content: it is that its absence is noticeable. A human who does not get their seven o’clock message knows something is wrong with the monitor itself.
Every five minutes
The hub pings a third-party service that nobody here controls, a free Healthchecks.io check or similar. If the hub dies, that service sends the email. Nothing should mark its own homework, and the ping is deliberately withheld whenever the sweep is not genuinely running, which is the failure that actually matters.
On the same screen as yours
Its own disk, memory and database go through the same checks and the same thresholds as anything else, and appear in the same list. It is a server too, and it is not exempt from its own rules.
What goes on your machines
It runs unattended on somebody else’s server for years. It must never fill a disk, never hang, never wedge a machine and never need attention. Everything about its design follows from that.
A check-in, in full
// figures, never content: see the privacy note below
{
"sentAt": "2026-09-08T14:05:00Z", // a fact, never a basis
"agentVersion": "1.0.0",
"uptimeSeconds": 862134,
"os": "Ubuntu 24.04.3 LTS",
"disks": [ { "mount": "/",
"totalBytes": 487880318976,
"freeBytes": 357411692544 } ],
"memory": { "totalBytes": 8232421376,
"availableBytes": 5911236608 },
"load": [ 0.05, 0.02, 0.00 ],
"services": [ { "name": "docker", "running": true } ],
"backup": { "lastSucceededAt": "2026-09-08T01:09:31Z",
"outcome": "ok" }
}
Installing it
# Linux, the token is baked in by the hub
curl -fsSL https://hub.redkite.info/install.sh \
| sudo bash -s -- --token rk_live_9f3a…
# The manual steps are published too, because piping a
# script into sudo bash is exactly what we would tell you
# never to do. It is short enough to read in full first.
A machine that can only manage “I am alive” is still worth watching. Every figure is optional, an agent that refuses to report because it cannot read one of them is an agent that reports nothing.
The line we will not cross
This is the most important decision in the whole product, and the one we are most often asked to bend. The reasoning is arithmetic rather than caution: an agent that can be told what to run turns one compromised hub into administrative access to every customer site at once. That is not a monitoring system. It is a botnet with a support contract.
| The reply may carry | Because the agent | Allowed |
|---|---|---|
| A suggested interval, in seconds | writes it to a config value, as a bounded integer | Yes |
| A replacement token | writes it to a config file, after checking its prefix and character set | Yes |
| A URL to fetch | would have to go and get something | No |
| A script, command, path or package version | would have to run something | No |
| A flag changing which figures are gathered | would be taking its behaviour from the network | No |
The test is simple: could a malicious reply from the hub cause the agent to execute code, read a file it does not already read, or contact a host other than the one in its local configuration? An integer and an opaque credential fail that test safely. Anything with a scheme, a slash or a shell metacharacter in it does not.
There is no auto-update either. A bad update lands on every machine simultaneously, and the agent’s entire value is that it is dull and always works. Updates are deliberate, staged, and pushed by an engineer.
Somebody will ask for remote access, and it is a genuinely useful feature. The answer is that it belongs in a separate tool with its own authentication, its own audit trail and its own explicit consent, not bolted onto the one piece of software allowed to run unattended on every machine you own.
Anything proposed for the payload is judged by one question: if this database leaked, what would the customer mind? Tokens are one per machine, hashed at rest, prefixed rk_live_ so one found in a log is recognisable, revocable without a site visit, and compared in fixed time.
Said here for the same reason it is said on screen
Every monitoring product has blind spots. With most of them, you find out about theirs from a customer. Here are ours, in writing, before you install anything.
Blind spot
PERC, SmartArray and LSI each need their own vendor tool, none of which can be assumed present and none of which this agent will shell out to. Where one is in use, the hub reports the blind spot rather than an absent array.
Blind spot
If the host owns the disk, the guest cannot see its health. We say so on the machine rather than leaving a gap that reads like a clean bill.
Blind spot
The installer says so at install time, and the hub keeps saying so afterwards until somebody fixes it.
Out of scope, deliberately
Is the website returning the right content, is the query slow, a different product, done well by other tools. Red Kite watches the machine underneath it. We also do not collect logs, and we do not poll switches by SNMP, because SNMP has to poll inwards, which is the one thing this design exists to avoid.
The beta
Red Kite is finished enough to be useful and not finished enough to sell. So it is free, and what we would like in return is to hear what it gets wrong on your machines rather than only on ours.
No account, no card, and no trial clock ticking down in the corner.
Please do not make this the only thing watching your machines yet. A monitoring system that is trusted and then misses something is worse than having none at all, that argument is most of this website, and it applies to us as much as to anybody else. Run Red Kite alongside whatever you already have, and tell us where the two disagree. That disagreement is exactly what we are looking for.
What we would like back
A figure it reported that was not true. An alert that never arrived. An alert that arrived and should not have. A machine it could not read at all. Those are worth more to us than praise, and there is no way to find them except on somebody else’s hardware.
What happens afterwards
We would rather say so now than surprise you later. If that changes we will say so on this page well before it does. Anything already installed keeps running: there is no licence check to fail, because there is no licence check, and we are not going to add one to software people installed while it was free.
The plain position
That is the ordinary footing for software given away, and on a monitoring tool this early it is also the honest one: we do not yet have enough machines behind us to promise it catches everything. What we can promise is that it will never quietly report a problem as fine.
Where to put the hub
You can install Red Kite anywhere that runs Docker: a machine in your own rack, or a five-pound virtual server somewhere else entirely. Both work. They are not equally good ideas, and it would be dishonest to pretend otherwise.
Recommended
The hub has to be outside the building. That is the entire premise of the product. A hub in your own server room shares the power, the broadband and the roof with the machines it is watching, so on the morning that actually matters, the monitor dies alongside them, and silence from a dead monitor looks exactly like a quiet night.
It also keeps the exposure off your network. Agents must be able to reach the hub, which means the hub is reachable from the internet. Host it internally and you are publishing a service from inside your own perimeter, with a route from that service to everything else on the LAN. Host it on its own small server elsewhere and the worst case is contained: it holds figures, and it holds no route back into any of your sites.
It is still yours. Your virtual machine, your database, your backups. Cloud here means “somewhere that is not your building”, not “somebody else’s service”.
Supported
Sometimes it is the only option, a policy that forbids outside hosting, a regulator who says so, or a site with no usable outbound route. It works, and we will help you do it.
Go in with both costs understood. You take on the shared failure domain, and you take on publishing an inbound service from your own network. If you must, put the hub in a DMZ with no path back to the LAN, on its own power and ideally its own line.
It is a much better fit for watching machines at other sites than for watching the ones in the same room. A hub in your own rack is an excellent way to watch twenty customer servers across the county. It is a poor way to watch the server standing next to it.
Whatever the licence, this is the bill you carry afterwards, and it goes to a hosting company, not to us. Eighty machines checking in every five minutes is about 23,000 rows a day, which is trivial for PostgreSQL on the smallest server you can rent.
The cost driver is not scale, it is text messages during a bad week. Push notifications are effectively free and are the right default for almost everything; keep SMS for critical and the bill stays in pennies.
| Small VPS (2 vCPU, 4 GB) | £4–6 |
|---|---|
| Domain | ~£1 |
| TLS certificate | Free |
| Transactional email | £0–12 |
| Push notifications | ~£4 once |
| Text messages | 2–4p each |
| External dead man’s switch | Free |
| All in | ~£6–20 |
Active development
Red Kite is worked on continuously. Owning a licence includes what comes next: new versions are yours to install when they suit you. Here is what is working today, what is being built, and what will never be built on purpose.
Suggest an improvement
Ideas go straight to the engineers who build it and are read by a person, not a form queue. There is no product committee here: if a suggestion is good and small, it often ships in the next version.
Send what you were trying to do and what got in the way, rather than a feature name; the underlying problem is usually more useful than the proposed solution. You will get an answer either way, including a plain no with the reason.
Some answers are permanent, and it is fairer to say so up front: anything that requires the agent to accept instructions, or the hub to connect into a monitored network, will be declined however it is framed. Those two lines are the product.
Put “improvement suggestion” in the subject and it goes to the technical side for review rather than into the sales pile.
Questions people actually ask
No, and deliberately not. An RMM exists to reach into machines and change them; Red Kite exists to notice when one stops talking. The two goals need opposite security postures, and combining them is how a monitoring agent becomes the most dangerous software on a customer’s network.
It sits happily alongside an RMM. Plenty of firms run both, and Red Kite keeps working on the day the RMM’s own agent stops checking in.
Nothing, while it is in beta. No subscription, no account with us, no card, and no trial clock ticking down. You install it on your own server and it is yours to run: if we vanished tomorrow your hub would carry on watching your machines, because nothing in the running system depends on us being there.
The only cost is hosting, and that goes to whoever rents you the server, roughly £6 to £20 a month all in, whether you are watching five machines or five hundred.
Possibly, and we would rather say so plainly now than let you find out later. Red Kite is in beta because it is not finished, and at some point it may well have a price on it. If that happens we will say so on this page well before it does.
What will not happen is anything you already have installed stopping. There is no licence check to fail, because there is no licence check, and we are not going to add one to software people installed while it was free. That is a promise about our own future conduct, which is the only kind worth writing down.
Write to contact@redkite.info and say roughly what you would like to watch, or fill in the beta form, which asks the same things. There is no checkout and no card wall, because there is nothing to buy.
What we would like back is what it gets wrong on your machines, a figure that was not true, an alert that did not arrive, a disk it could not read. We have run it on our own hardware for as long as that usefully teaches us anything; everything we learn from here has to come from somebody else’s.
No. Watch one machine or eighty. There is no per-agent fee, no per-endpoint fee, no per-technician licence and no sensors to count: during the beta or after it. Adding a machine is issuing it a token.
This matters more than it sounds. Per-device pricing quietly teaches people to leave the cheap machines unwatched, and the forgotten NAS in the corner is very often the one that dies.
The cloud, almost always, and by “cloud” we only mean “somewhere that is not your building”. A hub in your own server room shares the power, the broadband and the roof with the machines it is watching, so the outage you most need to hear about is the one that takes the monitor down too.
There is a second reason. Agents have to reach the hub, so the hub is reachable from the internet. Running it internally means publishing a service from inside your own perimeter; running it on a small server elsewhere keeps that exposure off your network entirely, and the machine holds figures rather than a route back into any of your sites.
Your own server room is supported and sometimes necessary. It is a much better fit for watching machines at other sites than the ones standing next to it.
The default is a check-in every five minutes with a fifteen-minute grace period, so an incident opens roughly twenty minutes after a machine goes quiet and the message leaves within the minute after that. Both figures are per machine: a critical server can be tighter, a laptop that sleeps can be much looser.
The grace period exists because networks hiccup and laptops sleep. A monitor that cries wolf at every dropped packet is one people mute, and a muted monitor is worth nothing.
Figures, never content. Byte totals, counters, up/down states, and how many of something there were. No file names, no user names, no process command lines, no log message text, no drive serial numbers and no traffic figures.
The published payload above is the whole of it, and anything new is judged by one question before it is added: if this database leaked, what would the customer mind?
Nothing breaks. The hub timestamps every check-in on its own clock, so a machine forty minutes out is still monitored correctly. Its clock travels as a fact, and we tell you about the drift: “this machine’s clock is 41 minutes behind ours, so its own timings may look odd”, because drift breaks certificate validation, log correlation and domain sign-in, in that order, weeks before anybody notices.
No. One outbound HTTPS connection on port 443, which every site already permits. No inbound rule, no port forward, no VPN, no static IP. It works behind NAT, behind CG-NAT, and through a corporate proxy.
Three things, all running at once: a daily digest whose absence tells you something is wrong; an external dead man’s switch run by a third party who emails you if your hub stops pinging it; and the hub monitoring itself as an ordinary machine on the same screen as yours.
The external ping is deliberately withheld whenever the sweep is not genuinely running, a hub that is up but not sweeping is the failure that actually matters, and it would otherwise look perfectly healthy.
Not yet, and we would rather be straight about it. The foundation is in from the first migration, every query is already filtered by the set of customers the caller may see, with staff holding a flag meaning “all”, because tenancy retro-fitted is a rewrite and designed in from the start costs nothing.
What has not been built is the customer-facing screens and per-customer alert routing. That work is in progress; see what we are building.
Wherever you put it. You run the hub, so the database is yours, a standard PostgreSQL database on your own server, which you back up and can read, export or take elsewhere at any time. We hold no copy of it and have no access to it.
If UK or EU residency matters to you, that is settled by where you rent the server, and it is entirely your choice. Nothing about Red Kite requires the data to leave the country it is created in.
Contact us
One server in a cupboard or eighty across a customer base: either is a short conversation. We will tell you honestly if it is not the right tool for you. If you would like to test it, the beta form is the quicker route: it asks the four things we would otherwise have to write back and ask you.
contact@redkite.infoRed Kite is built by a small team of systems engineers, and it runs on our own machines before it runs on anybody else’s. Who builds this.