Observability Case Study

Monitoring andObservability Platform

A central platform providing infrastructure, container, application and security visibility through metrics, logs, dashboards and operational alerting.

  • Prometheus
  • Grafana
  • Loki
  • Promtail
  • cAdvisor

Executive Summary

Turning technical telemetry into operational understanding.

The Monitoring and Observability Platform provides a central view of the engineering environment, including Linux hosts, Docker containers, security services and specialist applications.

The project began with a Grafana proof of concept intended to demonstrate how improved visibility could support technical teams. It developed into a broader platform using Prometheus, exporters, Loki, dashboards and alerting.

The platform is now used to understand performance, identify failures, investigate incidents, observe attack activity and provide live evidence of the environment in operation.

The Challenge

Move from reactive support to proactive visibility.

Without centralised monitoring, technical teams often discover service problems through user reports, manual checks or isolated application logs.

The challenge was to create monitoring that answered useful operational questions, supported troubleshooting and communicated service conditions clearly without overwhelming users with data.

Project Objectives

Monitoring designed around decisions and outcomes.

01

Create central visibility

Bring infrastructure, containers, applications and security telemetry into a consistent operational view.

02

Move beyond reactive support

Identify service-health issues and trends before they become significant incidents.

03

Support technical decisions

Provide reliable data that can be used during troubleshooting, capacity planning and service improvement.

04

Standardise monitoring

Use common metrics, dashboards and alerting patterns across multiple services.

05

Improve communication

Translate complex technical data into views that are understandable to support teams and stakeholders.

06

Provide live evidence

Publish sanitised dashboards demonstrating that the platform is actively monitored and operated.

Delivery Approach

From proof of concept to operational platform.

The solution was developed incrementally so that each stage demonstrated value before the next capability was introduced.

01

Identify the visibility gap

Recognised that operational teams required clearer, centralised visibility of service behaviour and infrastructure health.

02

Build the proof of concept

Used Grafana to demonstrate how technical data could be translated into useful operational dashboards.

03

Introduce metrics collection

Deployed Prometheus and exporters to collect consistent infrastructure, container and application metrics.

04

Develop dashboards

Built focused dashboards for platform health, security activity, container resources and specialist workloads.

05

Add operational alerting

Configured alerts for failed targets, exporter outages and unusual operational conditions.

06

Publish selected evidence

Created sanitised, read-only Grafana dashboards for inclusion in the public portfolio.

Data Sources

Telemetry from across the engineering platform.

Host Metrics

Node Exporter

Provides CPU, memory, disk, filesystem, network and operating-system telemetry.

Container Metrics

cAdvisor

Provides container resource usage, lifecycle and performance information.

Security Metrics

CrowdSec Exporters

Expose CrowdSec decisions, alerts and Linux firewall blocking activity.

Application Metrics

BirdNET Exporter

Provides detection and operational data from the BirdNET specialist workload.

Metrics Platform

Prometheus

Scrapes, stores and queries time-series metrics across the engineering environment.

Log Platform

Loki and Promtail

Collect and centralise host, application, container and reverse-proxy logs.

Monitoring Architecture

Collection, storage, visualisation and response.

01Infrastructure and ApplicationsHosts, containers, security tools and workloads
02Exporters and PromtailExpose metrics and forward logs
03Prometheus and LokiStore time-series metrics and logs
04GrafanaDashboards, queries and operational views
05Alerts and ResponseNotification, investigation and corrective action

Operational Dashboards

Views designed for different operational questions.

Security Monitoring

Purpose

Shows threat-intelligence decisions, local detections, blocked packets and malicious traffic.

Audience

Security, infrastructure and operational support.

Platform Health

Purpose

Provides visibility of host availability, CPU, memory, storage, containers and monitoring targets.

Audience

Infrastructure and platform operations.

BirdNET Activity

Purpose

Demonstrates application-specific monitoring for a real audio-processing workload.

Audience

Application support and technical demonstration.

Service Availability

Purpose

Tracks the health and responsiveness of key applications and supporting services.

Audience

Operational support and service management.

Project Delivery

Management principles applied to technical improvement.

The project demonstrates responsibilities aligned with project management as well as monitoring engineering.

01

Requirements definition

Identified the operational questions dashboards needed to answer before selecting panels and metrics.

02

Stakeholder communication

Presented monitoring information in a format that could be understood by engineers, support teams and non-specialists.

03

Incremental delivery

Started with a proof of concept and added metrics, dashboards and alerts in controlled phases.

04

Risk management

Considered alert fatigue, public-data exposure, target failure and monitoring-platform availability.

05

Operational handover

Documented dashboards, data sources, alert logic and troubleshooting procedures.

06

Continuous improvement

Refined queries and dashboards based on actual platform behaviour and support requirements.

Engineering Challenges

Balancing visibility, reliability and usability.

Selecting useful metrics

Problem

Collecting every available metric would create noise without improving operational understanding.

Response

Focused dashboards on service health, failure conditions, resource trends and security activity.

Outcome

Dashboards remain useful during troubleshooting and operational review.

Exporter integration

Problem

Different services expose metrics in different formats, and some security information was not directly available.

Response

Used native exporters where possible and developed custom exporters for CrowdSec firewall activity and BirdNET.

Outcome

Prometheus now collects consistent telemetry across infrastructure, containers, applications and security controls.

Alert noise

Problem

Poorly designed alerts can create repeated notifications without providing actionable information.

Response

Used reduction and threshold expressions, sensible evaluation periods and focused failure conditions.

Outcome

Alerts identify meaningful service-health conditions while reducing unnecessary noise.

Public dashboard security

Problem

Live dashboards needed to demonstrate operational evidence without exposing internal addresses, paths or administrative access.

Response

Created sanitised dashboard copies, used read-only external sharing and configured path-specific authentication bypass rules.

Outcome

Visitors can view selected telemetry while the main Grafana environment remains protected by MFA.

Operational Outcomes

Monitoring that supports real platform operation.

24/7

Operational visibility

Continuous collection and visualisation of platform telemetry.

Multi-source

Central monitoring

Infrastructure, container, security and application metrics in one system.

Proactive

Alerting

Notification of failures and unusual conditions before manual discovery.

Public

Live evidence

Sanitised dashboards demonstrating a real operating environment.

Skills Demonstrated

Technical monitoring and structured service improvement.

01Monitoring strategy
02Requirements analysis
03Prometheus
04Grafana
05Loki and Promtail
06Metrics and exporters
07Dashboard design
08Alert engineering
09Operational support
10Stakeholder communication
11Risk management
12Service improvement

Lessons Learned

Useful monitoring starts with the questions being asked.

Monitoring platforms should be designed around operational decisions rather than the number of metrics available.

A successful proof of concept needs to communicate value clearly, not only demonstrate technical capability.

Alerts must be actionable. Repeated or unclear notifications reduce confidence in the monitoring platform.

Public visibility can be provided safely when dashboards are deliberately sanitised and access rules are narrowly scoped.

Live Evidence

Explore the monitoring platform in operation.

Public dashboards provide read-only operational evidence while sensitive infrastructure details remain protected.