Terra Technologies Let’s talk

Consulting · ITOM

ITOM

You find out about outages from users. Alerts fire into a channel nobody owns. Availability is asserted rather than evidenced. [Draft copy]

The problem

Monitoring is not the same thing as knowing.

Monitoring accumulated the way the estate did: a tool per platform, thresholds set by whoever installed it, and an alert channel that everyone has muted. The signal exists. The ownership does not.

So the first notice of an outage is a user, capacity planning is a conversation rather than a number, and when a customer asks for last year’s availability against the SLA, the honest answer is that it would take a fortnight to reconstruct.

What enterprise-grade looks like here

The difference is that an alert means somebody moves.

Enterprise-grade operations means discovery that keeps the inventory honest without anyone maintaining a spreadsheet, event correlation so that one failure raises one alert rather than forty, and every alert routed to a named owner with a defined response.

Availability is then a report you run, not a reconstruction you commission — which is the difference between answering a due-diligence questionnaire in an afternoon and in a month.

This framing repeats on every capability page. It is the differentiator, and it reads consistently across all of them (§9 T3).

What we do

Four moves, in this order.

  1. Discover the estate and keep it discovered

    Automated inventory of what is actually running, reconciled against what you believe is running. The gap is usually the interesting part.

  2. Design monitoring around services, not servers

    Thresholds that reflect what the business notices, and correlation so a single fault does not produce a wall of noise.

  3. Route every alert to a named owner

    On-call rotation, escalation, and a response expectation per service. An alert with no owner is deleted, not left firing.

  4. Turn operations into evidence

    Availability, incident volume and mean time to restore, reported from the same data operations already produces.

[Draft copy]

What you end up owning

Documents and configuration, not a dependency.

Deliverables and the team that owns each one after handover
DeliverableOwned by
Automated estate inventoryOperations
Service-based monitoring and correlation rulesOperations
Alert routing with named owners and escalationOperations
Availability and MTTR reportingIT leadership
Capacity and utilisation baselineIT leadership and finance

[Draft copy]

Tell us what you are being asked to prove.

Tell us what you are being asked to evidence about availability, and we will tell you whether the data already exists.

Start a conversation