Platform Engineering Manager (Ref 23939AL)

Jobline Resources · Singapore

Sector
Data & Analytics
Function
Product & Engineering
Level
Mid-Level
Employment type
Full Time
Posted
2026-09-02
Source
mycareersfuture

Responsibilities: ·       Deploy and Operate Scalable Platforms - Implement, support, and optimize scalable, resilient, and highly available cloud-native, hybrid-cloud, and on-premises platform solutions that support enterprise applications and technology services. Collaborate with development and architecture teams to ensure platform reliability, performance, security, and operational efficiency.·       Maintain Enterprise Platform Standards and Best Practices - Adhere to established architectural guidelines, design principles, and technology standards for cloud platforms, microservices, container orchestration, messaging systems, databases, and integration services. Contribute to continuous improvements that enhance platform stability, scalability, and maintainability.·       Implement and Support Container Platforms - Deploy, administer, and continuously improve Kubernetes-based platforms across cloud and on-premises environments, including AWS EKS, ECS, and enterprise data centres. Manage Helm-based application deployments, platform upgrades, capacity planning, infrastructure automation, and troubleshooting to ensure highly available and resilient services.·       Implement Kubernetes Security, Governance and Compliance Controls - Design, implement, and maintain Kubernetes security policies using tools such as Kyverno and container image scanning solutions. Enforce RBAC, Network Policies, Pod Security Standards, admission controls, secrets management, workload isolation, and security baselines. Continuously monitor platform compliance, perform vulnerability assessments, support security audits, and ensure adherence to enterprise security, governance, and regulatory requirements across all Kubernetes environment·       Manage Event-Driven and Messaging Platforms - Deploy, configure, and support messaging and streaming platforms including Apache Kafka (KRaft) and RabbitMQ. Monitor platform health, optimize performance, manage cluster operations, and support event-driven architectures that enable scalable and loosely coupled distributed systems.·       Build and Maintain CI/CD Pipelines – Design, implement, and support automated CI/CD pipelines using Jenkins, ArgoCD, GitLab, and related DevOps tools. Ensure secure, reliable, and efficient software delivery through automation, deployment validation, release orchestration, and adherence to DevSecOps and governance requirements.·       Implement Secure Platform and Identity Services - Configure, deploy, and support enterprise authentication and authorization solutions using OAuth 2.0, JWT, Single Sign-On (SSO), and Identity & Access Management (IAM) platforms including Keycloak, Oracle IDMS (Oracle Access Manager), and Kerberos-based Windows Native Authentication (WNA). Integrate Active Directory (AD/LDAP) services and administer user federation, RBAC, access policies, and security controls to ensure compliance with enterprise security standards.·       Develop and Maintain Infrastructure as Code (IaC) - Implement, manage, and automate infrastructure provisioning and configuration management using Terraform, Ansible, and similar technologies. Ensure consistency, repeatability, scalability, and operational excellence across development, testing, and production environments.·       Administer Enterprise Database Platforms - Deploy, manage, and optimize PostgreSQL (Patroni clustering), MSSQL, MySQL, and MongoDB environments. Support database availability, performance tuning, backup and recovery processes, replication, disaster recovery, and operational monitoring to ensure reliable and resilient database services.·       Monitor, Troubleshoot, and Improve Platform Reliability - Implement monitoring, logging, alerting, and observability solutions across infrastructure and application platforms using Grafana, Prometheus, Manage Engine etc. Proactively identify performance bottlenecks, resolve incidents, conduct root cause analysis, and implement preventive measures to improve system reliability and operational efficiency. ·       Drive Automation and Continuous Improvement - Identify opportunities to eliminate manual processes through automation and scripting using shell scripts, ansible etc. Contribute to operational excellence initiatives, platform modernization efforts, and continuous improvement of deployment, monitoring, security, and infrastructure management practices.Requirement: ·       Bachelor of Science with 3-5 years Proven hands-on experience deploying, administering, and troubleshooting Kubernetes-based platforms using Docker and Helm, supporting large-scale microservices and containerized environments. ·       Strong expertise in cloud, hybrid-cloud, and on-premises infrastructure environments, with hands-on experience managing AWS services and on-premises virtualization platforms such as VMware ESXi and vSphere. ·       Deep understanding of DevOps practices and CI/CD implementation, with practical experience building, maintaining, and optimizing automated deployment pipelines using Jenkins, ArgoCD, GitLab CI/CD, or equivalent platforms. ·       Hands-on experience deploying, administering, and supporting event-driven and messaging platforms, including Apache Kafka and RabbitMQ, for reliable, scalable, and high-performance distributed systems.·       Strong background in database administration and platform operations with hands-on experience deploying, managing, and optimizing PostgreSQL (including high-availability clustering), MySQL, Microsoft SQL Server, Cassandra, Elasticsearch, and Trino. Experience supporting scalable, resilient, and high-performance data platforms for transactional systems, enterprise search, analytics, reporting, and business intelligence solutions, including Microsoft Power BI.·       Expertise in identity and access management technologies, including OAuth 2.0, SAML, JWT, and enterprise Single Sign-On (SSO) implementations using Keycloak, Kerberos-based Windows Native Authentication (WNA), and Oracle IDMS (Oracle Access Manager).·       Proficiency in Infrastructure as Code (IaC) and configuration management, utilizing Terraform, Ansible, Chef, or equivalent automation tools to provision, configure, and maintain infrastructure consistently across environments.·       Strong automation and scripting skills using Python, Shell, Ansible, terraform etc., to automate operational processes, platform management, monitoring, deployments, and system integrations.·       Proven ability to manage platform delivery and operational support, working closely with development, infrastructure, security, and architecture teams to implement, maintain, and continuously improve enterprise platforms and services.·       Strong experience in technical troubleshooting, incident management, and root cause analysis, with the ability to diagnose complex platforms, infrastructure, and application issues and implement sustainable solutions.·       Experience implementing and supporting monitoring, logging, and observability solutions using enterprise monitoring tools to ensure platform availability, performance, reliability, and operational excellence. ·       Ability to collaborate effectively in cross-functional engineering teams, participating in technical discussions, infrastructure reviews, deployment planning, and operational readiness activities to ensure successful solution delivery and support.Shortlisted candidates will be offered a 1 Year agency contract employment. License No : 12C6060

Apply on mycareersfuture →
Data & Analytics MASSIVE Distributed Platforms Scalability Design Practices Database Systems Platform Management Container Orchestration