NeverHard

Principal Site Reliability Engineer – Performance (A&D, Ultra HA/Exadata) at IFS — NeverHard

Principal Site Reliability Engineer – Performance (A&D, Ultra HA/Exadata) at IFS in Ottawa, Ottawa region. Skills: Architecture, Automation, Cloud-Native, Database Performance Tuning, Exadata. Apply on NeverHard.

Company
IFS
Location
Ottawa, Ottawa region
Type
full_time

Required skills:

Job Description Job Description The role of IFS Principal Site Reliability Engineer – Performance Engineering (Principal-SRE) exists within the Unified Support organization/division and serves as a technical and strategic leader for Site Reliability Engineering practices. By being part of the shift operation and broader SRE leadership, a Principal-SRE will drive 24x7x365 support excellence to IFS customers across the globe within the Aerospace & Defense (A&D) industry vertical while establishing and elevating SRE practices across the organization. This role sits within a dedicated performance engineering team supporting our Ultra High Availability (Ultra HA) on Exadata customers, where sustained database and application performance is a contractual commitment rather than a best-effort goal. As Principal SRE – Performance Engineering, you will be responsible for architecting and evolving the reliability frameworks, automation platforms, and operational strategies that underpin IFS's cloud-native products and infrastructure. This role involves setting technical direction, leading cross-functional teams, and partnering with R&D, Product, and Operations leadership to ensure that reliability is embedded at every level of service delivery. You will be an influential technical leader who drives organizational transformation through SRE principles, automation, and continuous improvement culture. Strategic Impact Areas Reliability Architecture & Strategy: Define and evolve the SRE operating model for the A&D vertical, including incident management, capacity planning, and disaster recovery strategies. Architect resilience patterns and frameworks for cloud-based applications supporting critical Aerospace & Defense operations. Establish Service Level Objectives (SLOs) and Error Budgets aligned with customer contractual requirements and business objectives. Drive adoption of chaos engineering and resilience testing across production systems. Technical Leadership & Platform Engineering: Lead the design and implementation of enterprise-scale monitoring, observability, and automated remediation platforms. Establish and maintain infrastructure-as-code standards and CI/CD pipelines for reliable deployments. Champion modern SRE tooling and methodologies across Unified Support and R&D organizations. Define technical architecture for Kubernetes, multi-cloud orchestration, and cloud-native application patterns. Architect solutions for performance optimization, cost efficiency, and security at scale (3,200+ concurrent users, 6,048+ database connections). Ultra HA & Exadata Performance Engineering: Own the performance and availability posture of Ultra HA environments running on Oracle Exadata, including proactive database health monitoring, workload and wait-event analysis, and end-to-end application performance management. Establish deep observability across the Oracle stack and the application tiers it serves, using Elastic, Grafana, and OpenTelemetry-based instrumentation to correlate database behaviour with user-facing latency. Lead performance troubleshooting for the most demanding customer workloads and drive tuning, capacity, and remediation decisions ahead of SLA impact. Organizational Transformation & Culture: Drive automation-first culture, eliminating toil and enabling teams to focus on high-value, creative work. Lead post-incident reviews and establish blameless culture focused on systemic improvement. Mentor and develop SRE team members, establishing career progression paths and technical mastery standards. Establish communities of practice for knowledge sharing across Unified Support and R&D organizations. Key Responsibilities Strategic Leadership & Architecture Set technical direction and establish long-term reliability strategies for IFS cloud products supporting the A&D industry Partner with R&D, Product, and Customer Success leadership to define reliability requirements and drive architectural decisions Design and oversee implementation of enterprise-scale observability, automation, and incident response platforms Lead architectural assessments and provide recommendations on infrastructure resilience, capacity planning, and disaster recovery Drive adoption of SRE best practices across product teams, establishing metrics and accountability for reliability Performance Engineering & Ultra HA Operations Own proactive performance monitoring for Ultra HA on Exadata customers — database health, workload profiles, wait events, and application response times — and act on trends before they breach SLA Lead deep performance investigations spanning the Oracle database, application servers, and infrastructure, producing evidence-based tuning and capacity recommendations Design and maintain the performance observability stack (Elastic, Grafana, OpenTelemetry/APM), including dashboards, baselines, and alert thresholds tied to customer SLOs Run performance baselining, load testing, and pre-release validation for Ultra HA environments, and quantify the impact of changes before they reach production Produce customer-facing performance reporting and participate in service reviews for Ultra HA accounts Incident Management & Operational Excellence Lead resolution of long-running, complex, and critical incidents affecting A&D customers; establish escalation frameworks and decision-making protocols Establish incident command systems, on-call rotations, and escalation procedures that balance response speed with team sustainability Drive post-incident review processes that focus on identifying systemic improvements rather than individual blame Establish incident trends analysis and drive prevention of recurrence through root cause elimination Automation & Continuous Improvement Champion the elimination of toil through intelligent automation, leveraging AI to reduce redundant, admin-heavy tasks Design and implement self-healing systems, automated remediation, and predictive alerting to minimize manual intervention Establish continuous improvement programs that measure, track, and drive down Mean Time To Detect (MTTD) and Mean Time To Recovery (MTTR) Drive infrastructure-as-code adoption, standardization, and versioning across cloud platforms (Azure, AWS, GCP) Establish runbook automation and GitOps practices for reliable, auditable deployments Documentation & Knowledge Management Establish knowledge management frameworks and internal KBAs that guide support operations for cloud-based applications Create and maintain architecture documentation, disaster recovery playbooks, and operational runbooks Lead documentation standards that ensure knowledge is accessible, accurate, and actionable for all support tiers Identify systemic knowledge gaps and drive closure through training, documentation, and process improvement Cross-Functional Collaboration Liaise with R&D and Unified Support Engineering to define automation requirements and platform capabilities for cloud applications Partner with customer success and professional services teams to translate customer reliability requirements into technical strategies Lead working groups and architecture reviews with multiple stakeholders to ensure reliability-first design decisions Establish feedback loops between support operations and product development to drive product reliability improvements Team Development & Mentorship Mentor and develop SRE team members, establishing technical mastery standards and career progression frameworks Lead technical training programs on cloud operations, Kubernetes, performance engineering, and SRE methodologies Foster a culture of learning, experimentation, and continuous improvement within the SRE organization Establish on-call culture that balances operational needs with team well-being and sustainable practices