AI Engineer (AWS, Splunk ITSI, Datadog, Python, New Relic, Azure, Dynatrace, ML, Service Now, Ansible, Rest API, ITIL)
Evyonic Solutions · Singapore
Responsibilities: Design, implement, and optimize enterprise observability solutions using Splunk ITSI. Architect and maintain observability solutions using Datadog for infrastructure, application, log, and service monitoring. Implement and optimize Dynatrace for application performance monitoring, infrastructure visibility, and service health analysis. Utilize New Relic for application monitoring, performance analysis, and proactive issue detection. Implement AppDynamics for application performance monitoring, transaction analysis, and application health management. Develop and enhance dashboards, KPIs, service views, alerts, and service-health monitoring for critical applications and infrastructure. Analyze logs, events, metrics, and telemetry to investigate incidents, identify anomalies, and support root-cause analysis. Implement intelligent event correlation, alert optimization, and noise-reduction techniques to improve operational efficiency. Apply AIOps, machine learning, and predictive analytics for anomaly detection, predictive monitoring, and proactive incident management. Integrate observability platforms with ServiceNow, CMDB, CI relationships, and service mapping to improve service visibility and dependency mapping. Design and implement observability solutions across AWS, Azure, and GCP cloud environments. Establish and monitor SLI/SLO metrics to improve service reliability and operational performance. Automate monitoring and operational processes using Python, Linux, Ansible, and REST APIs. Collaborate with infrastructure, application, cloud, and operations teams to improve service reliability and observability standards. Requirements: 15+ years of experience in AIOps, observability, infrastructure monitoring, systems engineering, or related IT operations roles. Bachelor’s degree in computer science, Engineering, Information Technology, or equivalent. Strong hands-on experience with Splunk ITSI and enterprise observability platforms. Hands-on experience with Datadog, Dynatrace, New Relic, and/or AppDynamics. Experience in ServiceNow, CMDB, CI relationships, and service mapping. Experience with AIOps, machine learning, anomaly detection, event correlation, and predictive analytics. Experience in Python, Linux, Ansible, and REST APIs for automation and integration. Experience working with AWS, Azure, and/or GCP cloud environments. Strong understanding of ITIL, SRE, incident management, and SLI/SLO concepts. Strong analytical and troubleshooting skills using logs, metrics, and events. Experience in designing and supporting enterprise-scale observability and monitoring platforms. Strong stakeholder management, technical leadership, and mentoring skills.