Create central visibility
Bring infrastructure, containers, applications and security telemetry into a consistent operational view.
Observability Case Study
A central platform providing infrastructure, container, application and security visibility through metrics, logs, dashboards and operational alerting.
Executive Summary
The Monitoring and Observability Platform provides a central view of the engineering environment, including Linux hosts, Docker containers, security services and specialist applications.
The project began with a Grafana proof of concept intended to demonstrate how improved visibility could support technical teams. It developed into a broader platform using Prometheus, exporters, Loki, dashboards and alerting.
The platform is now used to understand performance, identify failures, investigate incidents, observe attack activity and provide live evidence of the environment in operation.
The Challenge
Without centralised monitoring, technical teams often discover service problems through user reports, manual checks or isolated application logs.
The challenge was to create monitoring that answered useful operational questions, supported troubleshooting and communicated service conditions clearly without overwhelming users with data.
Project Objectives
Bring infrastructure, containers, applications and security telemetry into a consistent operational view.
Identify service-health issues and trends before they become significant incidents.
Provide reliable data that can be used during troubleshooting, capacity planning and service improvement.
Use common metrics, dashboards and alerting patterns across multiple services.
Translate complex technical data into views that are understandable to support teams and stakeholders.
Publish sanitised dashboards demonstrating that the platform is actively monitored and operated.
Delivery Approach
The solution was developed incrementally so that each stage demonstrated value before the next capability was introduced.
Recognised that operational teams required clearer, centralised visibility of service behaviour and infrastructure health.
Used Grafana to demonstrate how technical data could be translated into useful operational dashboards.
Deployed Prometheus and exporters to collect consistent infrastructure, container and application metrics.
Built focused dashboards for platform health, security activity, container resources and specialist workloads.
Configured alerts for failed targets, exporter outages and unusual operational conditions.
Created sanitised, read-only Grafana dashboards for inclusion in the public portfolio.
Data Sources
Host Metrics
Container Metrics
Security Metrics
Application Metrics
Metrics Platform
Log Platform
Monitoring Architecture
Operational Dashboards
Shows threat-intelligence decisions, local detections, blocked packets and malicious traffic.
Security, infrastructure and operational support.
Provides visibility of host availability, CPU, memory, storage, containers and monitoring targets.
Infrastructure and platform operations.
Demonstrates application-specific monitoring for a real audio-processing workload.
Application support and technical demonstration.
Tracks the health and responsiveness of key applications and supporting services.
Operational support and service management.
Project Delivery
The project demonstrates responsibilities aligned with project management as well as monitoring engineering.
Identified the operational questions dashboards needed to answer before selecting panels and metrics.
Presented monitoring information in a format that could be understood by engineers, support teams and non-specialists.
Started with a proof of concept and added metrics, dashboards and alerts in controlled phases.
Considered alert fatigue, public-data exposure, target failure and monitoring-platform availability.
Documented dashboards, data sources, alert logic and troubleshooting procedures.
Refined queries and dashboards based on actual platform behaviour and support requirements.
Engineering Challenges
Collecting every available metric would create noise without improving operational understanding.
Focused dashboards on service health, failure conditions, resource trends and security activity.
Dashboards remain useful during troubleshooting and operational review.
Different services expose metrics in different formats, and some security information was not directly available.
Used native exporters where possible and developed custom exporters for CrowdSec firewall activity and BirdNET.
Prometheus now collects consistent telemetry across infrastructure, containers, applications and security controls.
Poorly designed alerts can create repeated notifications without providing actionable information.
Used reduction and threshold expressions, sensible evaluation periods and focused failure conditions.
Alerts identify meaningful service-health conditions while reducing unnecessary noise.
Live dashboards needed to demonstrate operational evidence without exposing internal addresses, paths or administrative access.
Created sanitised dashboard copies, used read-only external sharing and configured path-specific authentication bypass rules.
Visitors can view selected telemetry while the main Grafana environment remains protected by MFA.
Operational Outcomes
Continuous collection and visualisation of platform telemetry.
Infrastructure, container, security and application metrics in one system.
Notification of failures and unusual conditions before manual discovery.
Sanitised dashboards demonstrating a real operating environment.
Skills Demonstrated
Lessons Learned
Monitoring platforms should be designed around operational decisions rather than the number of metrics available.
A successful proof of concept needs to communicate value clearly, not only demonstrate technical capability.
Alerts must be actionable. Repeated or unclear notifications reduce confidence in the monitoring platform.
Public visibility can be provided safely when dashboards are deliberately sanitised and access rules are narrowly scoped.
Live Evidence
Public dashboards provide read-only operational evidence while sensitive infrastructure details remain protected.