Own the build, configuration, and lifecycle management of Linux server fleets supporting colo trading (provisioning, patching, hardening, and performance tuning).
Engineer and maintain automation for OS deployment, configuration management, and continuous compliance to reduce manual touch.
Partner with network, trading technology, and data center teams to optimize latency, throughput, and stability; include kernel/IRQ/CPU tuning as needed.
Leverage AI-enabled reliability workflows for incident analysis and design, ensuring validation and data security.
Operate and enhance observability across the stack (metrics, logs, alerts) and translate signals into runbooks and service improvements.
Lead incident response for Linux/compute events; rapid triage, mitigation, root-cause analysis, and preventive actions.