Senior AI Platform Engineer

Bitdeer AI · Singapore

Sector
AI
Function
Product & Engineering
Level
Mid-Level
Employment type
Full Time
Posted
2026-10-08
Source
mycareersfuture

About Bitdeer:Bitdeer Technologies Group (Nasdaq: BTDR) is a world-leading technology company for Bitcoin mining. Bitdeer is committed to providing comprehensive computing solutions for its customers. The Company handles complex processes involved in computing such as equipment procurement, transport logistics, datacenter design and construction, equipment management, and daily operations. The Company also offers advanced cloud capabilities to customers with high demand for artificial intelligence. Headquartered in Singapore, Bitdeer has deployed datacenters in the United States, Norway, and Bhutan.Job Summary:We are seeking a Senior AI Platform Engineer to keep the MaaS inference platform reliable under real customer traffic. This role owns production stability across gateway, inference-proxy, control plane, model deployments, GPU clusters, and multi-region operations.What you will be responsible for:Operate and harden Kubernetes-based MaaS production environments across CPU platform nodes, edge ingress, and regional GPU tiers.Define and ownSLOs, alerting, dashboards, run books, and incident response for API availability, latency, error rate, capacity, and GPU health.Improve rollout safety for model/runtime/platform changes using canaries, fallback, health-aware routing, maintenance mode, and fast rollback.Drive capacity planning for GPU utilization, burst traffic, quota/rate limits, cross-region latency, and customer growth.Automate repetitive operations through Helm, Argo CD, operators, scripts, and self-healing workflows.Partner with runtime and performance engineers to debug incidents from public API edge to model worker.How you will stand out:6+ years in SRE, platform engineering, or infrastructure engineering for production cloud services.Deep Kubernetes experience, including Helm, Argo CD/GitOps, CNI/ingress, secrets, storage, and workload scheduling. Hands-on experience operating GPU, AI infrastructure, or HPC workloads is strongly preferred.Strong observability skills with Prometheus/VictoriaMetrics, OpenTelemetry, logs, traces, and incident diagnosis.Comfortable with Go, Python, Bash, Linux networking, and production automation.Proven ability to design reliable systems with clear SLOs, operational ownership, and post-incident follow-through.What you will experience working with us:A culture that values authenticity and diversity of thoughts and backgrounds;An inclusive and respectable environment with open workspaces and exciting start-up spirit;Fast-growing company with the chance to network with industrial pioneers and enthusiasts;Ability to contribute directly and make an impact on the future of the digital asset industry;Involvement in new projects, developing processes/systems;Personal accountability, autonomy, fast growth, and learning opportunities;Attractive welfare benefits and developmental opportunities such as training and mentoring.--------------------------------------------------------------------Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, colour, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.

Apply on mycareersfuture →
AI Software Debugging Dashboards systems reliability HPC Kubernetes Argo CD API Management