Observability Architect - FTC
Job in
Location, Tucker County, West Virginia, USA
Listed on 2026-07-18
Listing for:
AND Digital Limited
Full Time, Contract
position Listed on 2026-07-18
Job specializations:
-
IT/Tech
Cybersecurity, Systems Engineer, SRE/Site Reliability, IT Support
Job Description & How to Apply Below
Location: Location
Observability Architect
12 Month FTC
About you:
- You care deeply about producing high-quality work that delivers real value
- You’re comfortable navigating ambiguity and solving complex problems collaboratively
- You bring strong expertise in your craft, alongside a willingness to keep learning
- You communicate clearly and build trust quickly with clients and teammates
- You’re pragmatic, adaptable and outcome-focused
- You enjoy sharing knowledge and helping others grow
- You value low-ego collaboration and enjoy working as part of multidisciplinary teams
Lead the assessment, design, and optimisation of the observability strategy for the co-location migration programme. Ensure logging, metrics, tracing, alerting, and operational dashboards provide comprehensive visibility across the new infrastructure and application estate, enabling the successful migration of the tightly-coupled monolithic platform with minimal operational risk. Identify gaps in the existing observability capability and recommend enhancements to tooling, processes, and architecture where required.
Key ResponsibilitiesObservability Assessment & Strategy
- Review the current observability architecture across infrastructure, networks, middleware, databases, and applications.
- Assess existing logging, metrics, distributed tracing, and monitoring capabilities to determine readiness for the co-location migration.
- Develop an observability strategy that supports both migration activities and long-term operational support.
- Recommend enhancements or platform uplifts where current tooling does not provide sufficient visibility or resilience.
- Analyse telemetry, monitoring data, dashboards, and operational trends from completed migration waves.
- Establish performance baselines for compute, storage, networking, application response times, and transaction throughput.
- Identify recurring operational issues and use historical insights to improve migration readiness.
- Define measurable service health indicators to compare pre- and post-migration performance.
- Design comprehensive monitoring for the tightly-coupled monolithic application estate, with particular emphasis on latency-sensitive interdependencies.
- Create real-time dashboards that provide operational visibility across infrastructure, middleware, databases, messaging, and application components.
- Ensure end-to-end transaction tracing is available to rapidly identify bottlenecks and service degradation.
- Validate monitoring coverage prior to each migration wave.
- Review and standardise centralised logging across all migrated environments.
- Ensure consistent log formats, metadata, correlation IDs, and traceability across systems.
- Validate log ingestion, retention policies, indexing, and search performance.
- Ensure operational teams can rapidly investigate incidents using correlated logs and distributed traces.
- Review and optimise alert thresholds to minimise both missed events and unnecessary alert noise.
- Implement intelligent alerting aligned to business services and critical customer journeys.
- Define migration-specific alerting for infrastructure failures, application degradation, latency increases, replication issues, and capacity constraints.
- Support operational readiness activities including rehearsals and production cutover monitoring.
- Ensure observability solutions meet financial services regulatory requirements for auditability, log retention, security, and data governance.
- Validate access controls and security monitoring for observability platforms.
- Support evidence gathering for internal governance, audit, and regulatory reviews.
- Evaluate the suitability of existing observability platforms and recommend improvements where required.
- Assess opportunities to improve automation, anomaly detection, service health monitoring, and predictive alerting.
- Define standards and best practices for observability across future migration phases.
- Work closely with Infrastructure Architects, Application Architects, Platform Engineering, Security,…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×