Observability Architect - 12 Month FTC
AND Digital
| Company | AND Digital |
| Category | Engineering |
| Location | London |
| Remote | On-site (inferred) |
| Employment | Full-time |
| Level | Mid |
| Salary | Not stated by the employer |
| Posted | 14 Jul 2026 |
| Last verified | 7 Aug 2026 |
| Source | Employer ATS (workable) |
Description
Observability Architect 12 Month FTC Who We Are We’re on a mission to close the world’s tech skills gap. We help organisations navigate the future of technology, combining human expertise, emerging tech and AI to deliver better outcomes, faster. Since 2014, we’ve worked side-by-side with clients to solve complex challenges, build high-performing teams, and create lasting capability. As technology continues to evolve, we believe the most successful organisations will be those that combine the best of both: human ingenuity AND intelligent technology. That belief is embedded in everything we do. We call it the genius of the AND: deep expertise AND practical delivery, innovation AND responsibility, ambitious work AND sustainable careers. Through our Guide, Build and Equip approach, we help organisations embrace change, deliver meaningful impact, and develop the skills they need to thrive in an increasingly agentic world. About you: You care deeply about producing high-quality work that delivers real value You’re comfortable navigating ambiguity and solving complex problems collaboratively You bring strong expertise in your craft, alongside a willingness to keep learning You communicate clearly and build trust quickly with clients and teammates You’re pragmatic, adaptable and outcome-focused You enjoy sharing knowledge and helping others grow You value low-ego collaboration and enjoy working as part of multidisciplinary teams Role Objective Lead the assessment, design, and optimisation of the observability strategy for the co-location migration programme. Ensure logging, metrics, tracing, alerting, and operational dashboards provide comprehensive visibility across the new infrastructure and application estate, enabling the successful migration of the tightly-coupled monolithic platform with minimal operational risk. Identify gaps in the existing observability capability and recommend enhancements to tooling, processes, and architecture where required. Key Responsibilities Observability Assessment & Strategy Review the current observability architecture across infrastructure, networks, middleware, databases, and applications. Assess existing logging, metrics, distributed tracing, and monitoring capabilities to determine readiness for the co-location migration. Develop an observability strategy that supports both migration activities and long-term operational support. Recommend enhancements or platform uplifts where current tooling does not provide sufficient visibility or resilience. Baseline Performance Analysis Analyse telemetry, monitoring data, dashboards, and operational trends from completed migration waves. Establish performance baselines for compute, storage, networking, application response times, and transaction throughput. Identify recurring operational issues and use historical insights to improve migration readiness. Define measurable service health indicators to compare pre- and post-migration performance. Monolithic Application Monitoring Design comprehensive monitoring for the tightly-coupled monolithic application estate, with particular emphasis on latency-sensitive interdependencies. Create real-time dashboards that provide operational visibility across infrastructure, middleware, databases, messaging, and application components. Ensure end-to-end transaction tracing is available to rapidly identify bottlenecks and service degradation. Validate monitoring coverage prior to each migration wave. Logging & Trace Management Review and standardise centralised logging across all migrated environments. Ensure consistent log formats, metadata, correlation IDs, and traceability across systems. Validate log ingestion, retention policies, indexing, and search performance. Ensure operational teams can rapidly investigate incidents using correlated logs and distributed traces. Alerting & Operational Re