Software Engineer ( Dynatrace, GenAI, Prometheus ,L2 Support, Splunk)

Neptunez Singapore · Singapore

Sector
AI
Function
Product & Engineering
Level
Mid-Level
Employment type
Contract
Posted
2026-07-22
Source
mycareersfuture

ResponsibilitiesLead the Site Reliability Engineering (SRE) function for business-critical production platforms, ensuring high availability, reliability, and operational excellence.Provide technical leadership for Production Support (L2/L3), driving timely resolution of production incidents and minimizing service disruptions.Lead major incident management, root cause analysis (RCA), and post-incident reviews to improve platform stability.Utilize Dynatrace and Splunk to monitor application and infrastructure health, analyze performance trends, and proactively identify issues.Troubleshoot complex application, infrastructure, and database issues across distributed environments.Optimize SQL queries and support performance tuning across Oracle, MySQL, and MS SQL Server databases.Collaborate with development, infrastructure, database, and DevOps teams to improve system reliability and operational efficiency.Drive automation initiatives to streamline operational processes and reduce manual effort.Define and maintain operational standards, monitoring strategies, runbooks, and best practices.Mentor and guide junior engineers, provide technical direction, and foster knowledge sharing across the team.Engage with stakeholders to provide technical updates, operational insights, and continuous improvement recommendations.Requirements10+ years of experience in Site Reliability Engineering (SRE), Production Support, or Application Support, including experience in a technical leadership role.Strong hands-on expertise in Dynatrace and Splunk for monitoring, observability, alerting, and troubleshooting.Proven experience leading Production Support (L2/L3), incident management, root cause analysis (RCA), and service restoration.Strong expertise in SQL, including query optimization, performance tuning, and database troubleshooting on Oracle, MySQL, and MS SQL Server.Experience supporting distributed systems and mission-critical production environments.Good understanding of Java applications and application troubleshooting.Hands-on experience with Linux administration and scripting using Shell, Python, Ansible, Autosys.Hands-on experience in Linux, Microservices, Kafka.Hands-on experience in Azure.Experience in ITIL processes, Service Management, CI/CD, and automation practices.Experience in banking, healthcare domain preferred.Proven ability to mentor engineers, lead technical discussions, and drive operational improvements.Excellent analytical, problem-solving, communication, stakeholder management, and leadership skills.

Apply on mycareersfuture →
AI Distributed Database Infrastructure Monitoring optimized SQL queries Splunk Java Applications coordinating developers Business Function Architecture