Operate and improve reliability, scalability, and performance of the Japan Sovereign Cloud platform for Oracle Cloud Infrastructure (OCI), collaborating with software engineering, cloud operations, and global OCI teams; apply software engineering principles to automate operations, resolve complex production issues, and enhance service resiliency.
Participate in a 24x7 shift rotation, understand and follow 24x7 workflows, alerts, incidents, escalation paths, and runbooks; partner with Japanese and international shift teams to capture recurring operational issues, improve alert actionability, maintain operational documentation, and contribute fixes through tooling, automation, and process improvements.
Maintain and improve runbooks and handoff quality; drive day-to-day sovereign cloud operations, identify recurring pain points, and implement practical operational improvements.
技術スタック
必須スキル
Linux (production environments)
Python
Reliability Engineering
歓迎スキル(該当する場合)
Java, Go, Shell (or other scripting/programming languages)
Cloud computing, networking, distributed systems
Automation technologies and related tooling
Experience with 24x7 on-call operational support and incident response
Ability to improve runbooks, alert guidance, and operational handoffs
キャリア成長観点
Sovereign Cloud領域の安定運用と大規模クラウドインフラの信頼性設計に深く携わる機会により、SRE領域での技術力と運用スキルが体系的に成長。