Design and implement automation to ensure availability, scalability, and operational excellence of OCI Japan Sovereign Cloud services.
Lead complex incident investigations, perform root-cause analysis, and drive reliability improvements; partner with development teams to improve operational readiness.
Own and prioritize the SRD operational improvement backlog based on shift feedback, incident reviews, alert quality reviews, and business reliability requirements.
Translate operational and business requirements into reliability plans; execute improvements via tooling, automation, runbooks, and process changes; coordinate cross-team.
Participate in 24x7 shift rotation; provide technical leadership during critical service events.
Mentor less experienced engineers; contribute to continuous improvement across the organization.
Collaborate with JP Sovereign Cloud and EU Sovereign Cloud teams to share practices and align improvements.
Balance business requirements, technical feasibility, and operational risk when planning reliability improvements.
技術スタック
必須スキル
Linux
Python
Reliability Engineering
歓迎スキル(該当する場合)
Proficiency in additional programming languages (Java, Go, C++ or similar)
Experience with cloud platforms, infrastructure automation, observability/monitoring, and incident response
Strong Linux systems administration, networking, storage, and performance optimization
Experience with root-cause analysis, incident management, and improving alert quality
On-call leadership and cross-team collaboration
Native-level Japanese and business-level English communication