Provide technical leadership for the reliability, availability, and operational strategy of OCI's Japan Sovereign Cloud platform.
Lead large-scale reliability initiatives, influence architecture decisions, develop automation frameworks, and drive operational excellence across multiple cloud services.
Define SRD reliability strategy, operational standards, and improvement roadmaps aligning business requirements with technical execution (Alert, Incident Response, Availability, and Reliability).
Align JP Sovereign Cloud with EU Sovereign Cloud and global OCI reliability teams; drive standardization to improve service resiliency.
Participate in 24x7 operational support; act as escalation point for high-severity incidents; sponsor durable improvements from shift operations; mentor senior engineers on Plan + Execution ownership.
技術スタック
必須スキル
Linux
Python
Reliability Engineering
8+ years in SRE/Cloud Infrastructure/Distributed Systems
Experience designing and operating highly available cloud platforms
Expert-level software development and automation (Java/Go/Python or similar)
Distributed systems architecture, networking, storage, observability, and service resiliency
Proven incident response leadership and cross-organizational reliability programs
Able to influence architecture and engineering standards across multiple teams
Native Japanese and business-level English
歓迎スキル(該当する場合)
Cross-sovereign collaboration experience (JP Sovereign Cloud and EU Sovereign Cloud)
Ability to translate business/operational requirements into prioritized reliability roadmaps with measurable outcomes
Mentoring and coaching of senior engineers on reliability practices