AI Ops Engineer
Itcan · Singapore
Data & AI Operations Engineer #AIDABe a Part of Something BIG! Apply AI, ML, and LLM-based tools to solve real-world operational and reliability challengesIdentify, prototype, and validate AI-driven improvements to day-to-day operations using clear, measurable metricsOperationalise successful AI use cases into reliable, scalable, and maintainable production capabilitiesImprove system observability, incident detection, root-cause analysis, and operational efficiency through automation and intelligenceSupport the end-to-end lifecycle of AIOps, MLOps, and LLMOps solutions, from experimentation to productionEnsure operational reliability, robustness, and maintainability of AI-enabled systemsCollaborate closely with senior engineers and operations teams to align solutions with business and operational needsContinuously evaluate emerging AI technologies and tools for practical applicability and impactEmbed best practices for monitoring, governance, and responsible use of AI in operational environmentsMake an Impact by:Explore and evaluate AI, ML, and LLM-based tools that can improve operational efficiency and visibilityBuild small proof-of-concepts or prototypes to validate ideas in operational contextsMeasure outcomes using simple, clearly defined metricsAnalyze logs, metrics, traces, and events to identify anomalies and automation opportunitiesAssist in deploying, monitoring, and maintaining ML or LLM-enabled solutionsAutomate repetitive operational tasks using scripts, workflows, or lightweight servicesWork with senior engineers to turn validated ideas into stable, scalable capabilitiesDocument solutions, trade-offs, and lessons learned for reuse by other teamsStay informed on relevant AI and open-source tooling, and suggest ideas worth testingSkills for Success: Bachelor’s or Master’s degree in Computer Science or a related field1–2 year experience:Python, SQL, Apache Spark, or similar languagesData Wrangling, Analysis and VisualisationContainers and deployment fundamentals (Docker; Kubernetes)Version Control (Git)Testing (pytest)MLOps/GenAI Ops toolsInternships, graduate projects, hackathons, open-source contributions, or relevant work experience in:Software developmentAutomationAI / ML / GenAI-related projectFamiliarity withSoftware engineering best practices, including OO/Functional design, reproducibility, and testingLogs, metrics, and monitoring conceptsML/GenAI tooling (LLM, RAG, LlamaGuard, Presidio, MLFlow, LangFuse etc)Monitoring and observability stacks (Prometheus, Grafana, OpenTelmetry)Hands-on experience with cloud ML services (Databricks, Azure ML, etc.)Hands-on, curious, and proactivePragmatic and focused on operational impactDisciplined about scaling what works and discarding what doesn’tCollaborative and able to communicate ideas clearlySelf-driven and proactive, comfortable working in a fast-paced environmentFamiliarity with AI and data development process in telco environment