Site Reliability Engineer (Dynatrace, GenAI, Prometheus, L2 Support, Splunk)
Neptunez Singapore · Singapore
ResponsibilitiesDesign, implement, and continuously improve Site Reliability Engineering (SRE) practices to ensure highly available, scalable, and resilient production systems.Provide L2 production support by troubleshooting application, infrastructure, and platform issues while ensuring compliance with SLA and SLO commitments.Monitor, maintain, and optimize Microsoft Azure environments, including Azure Kubernetes Service (AKS), Azure Monitor, and Log Analytics.Support and troubleshoot Java and Spring Boot applications by analyzing logs, JVM performance, REST API issues, configuration problems, and application performance bottlenecks.Build, configure, and maintain observability solutions using AppDynamics, Dynatrace, Splunk, Grafana, and Prometheus to improve system visibility and proactive monitoring.Develop, enhance, and maintain CI/CD pipelines using Jenkins, Bitbucket, and Azure DevOps to support reliable and automated application deployments.Deploy, manage, and troubleshoot containerized applications running on Kubernetes clusters.Participate in incident management, on-call rotations, root cause analysis (RCA), and post-incident reviews to improve service reliability and reduce recurring issues.Collaborate with development, infrastructure, and DevOps teams to improve application performance, deployment processes, and operational stability.Utilize GenAI capabilities to improve operational efficiency through log analysis, incident summarization, automation, and troubleshooting assistance.Create and maintain operational documentation, monitoring dashboards, runbooks, and standard operating procedures.Manage incidents, change requests, and operational activities using Jira while ensuring timely resolution and effective communication with stakeholders.RequirementsBachelor's degree in Computer Science, Information Technology, Engineering, or a related field.8+ years of experience in Site Reliability Engineering (SRE), Production Support, or DevOps environments.Strong hands-on experience with Microsoft Azure, including AKS, Azure Monitor, and Log Analytics.Experience supporting Java and Spring Boot applications in production environments.Experience in Java application troubleshooting, application logs.Hand on experience in Microservices, RestAPIHands-on experience with observability and monitoring tools including AppDynamics, Dynatrace, Kibana , Splunk, Grafana, and Prometheus.Experience building and maintaining CI/CD pipelines using Jenkins, Bitbucket, and Azure DevOps.Hands on experience in managing Kubernetes.Experience providing L2 production support, incident management, root cause analysis, and problem management.Experience in Jira and change management.Exposure to GenAI tools and AI-driven operational automation is an added advantage.Strong analytical, troubleshooting, communication, and stakeholder management skills.Ability to work in a fast-paced production environment and participate in on-call support rotations.