AI Engineering Lead
Espire Infolabs Singapore · Singapore
The AI Engineering Lead is the technical authority of the AI Platform Engineering. This role owns the end-to-end engineering of all AI solutions — from solution design alongside the Enterprise and Cloud Architects through to production delivery — and is accountable for the technical quality of everything builds or procures. The role is responsible for defining designs, builds, tests, deploys, governs, and operates AI solutions on AWS. This includes Generative AI, Agentic AI, Retrieval-Augmented Generation, document intelligence, LLM integration, AI workflow orchestration, MLOps / LLMOps, and reusable AI engineering standards. The AI Engineering Lead works closely with Enterprise Architects, Cloud Architects, Cloud Platform Engineers, Data Science & Analytics teams, AI Governance, and delivery teams to ensure AI solutions are secure, scalable, reusable, cost-aware, and aligned to enterprise technology standards. This is a hands-on leadership role. The AI Engineering Lead is expected to guide solution design, review technical implementation, lead the engineering team, oversee vendor-delivered AI solutions, and establish production-grade engineering discipline across the AI Platform Engineering function. KEYRESPONSIBILITIES End-to-End Solution Design Own the technical solution design for all AI use cases — from intake through to production — producing architecture artefacts, design decisions and technical specifications Work alongside the Enterprise and Cloud Architects on end-to-end solution design — ensuring AI solutions are enterprise-compliant, secure and platform-aligned from day one Produce and review architecture artefacts, technical designs, integration patterns, engineering standards and implementation plans. Provide technical input into build / buy / partner decisions by assessing feasibility, complexity, delivery risk, maintainability, cost, and long-term platform fit. Ensure AI solutions are designed for reuse, operational support, observability, governance and future extensibility. AI Platform& Services — Design & Oversight Own the AI engineering layer on AWS — defining which native services are used for each use case and how they connect to the enterprise stack Lead technical design across the core AWS AI stack: Bedrock (LLM orchestration, agents, knowledge bases, guardrails), Agent Core, Sage Maker (ML model deployment, pipelines, MLOps), Textract/Comprehend (document intelligence), OpenSearch (vector search, RAG retrieval), Step Functions (pipeline orchestration) Work with the Cloud Platform Engineer — who owns environment provisioning and infrastructure setup — by providing clear technical requirements per use case so they can configure AWS environments correctly Own the LLMOps and MLOps architecture — model versioning, deployment, monitoring, drift detection, retraining triggers and production observability on AWS Oversee AWS AI service cost architecture — working with the AI Champion and Cloud Platform Engineer on FinOps tagging, per-use-case cost allocation and spend optimisation Production Engineering & Standards Ensure all AI solutions are built to production standards — with appropriate testing, security controls, PII handling, monitoring, logging, resilience, error handling and operational runbooks Own the AI Platform Engineering reuse library — ensuring components, patterns, prompt templates and AWS service wrappers built for one use case are packaged and available for reuse Lead technical production readiness reviews before any AI solution goes live — covering model risk, security, performance, cost and rollback plan Define standards for CI/CD, automated testing, release management, observability, incident response and operational support for AI workloads. Establish and maintain AI engineering runbooks, support models, incident playbooks, and operational standards for live AI solutions. Work with AI Governance and Risk teams to provide required technical artefacts such as architecture documents, model cards, risk mitigations, control evidence, and production readiness sign-offs. Ensure governance requirements are embedded into delivery processes without creating unnecessary delivery friction. Team Leadership Lead the AI Platform Engineering team across AI engineering, AI solution development, and AI testing disciplines. Provide technical direction, coaching, and development support to AI Engineers, AI Engineer Associates and AI Test Engineers. Set engineering culture and delivery discipline, including code review standards, documentation expectations, test coverage, secure coding, reuse, and operational accountability. Run technical design reviews, engineering forums, sprint planning discussions, and implementation reviews in collaboration with delivery and product stakeholders. Support team capability building across AWS AI services, GenAI engineering, Agentic AI, LLMOps, MLOps, and production AI engineering. Build a strong engineering environment that encourages ownership, learning, constructive challenge and quality delivery. DSA& Architecture Collaboration Define and own the technical handshake with the DSA team — clarifying where DSA's model development work ends and the AI Platform Engineering’s production engineering begins on each use case Engage with the Enterprise and Cloud Architects as a peer — contributing to architecture reviews, raising AI-specific platform requirements, ensuring solutions conform to enterprise standards Represent the technical capability in architecture design authority forums and technical governance reviews Vendor Technical Oversight Act as the technical counterpart for vendor-delivered or partner-delivered AI solutions. Review vendor architecture proposals, solution designs, integration approaches, security controls and deployment plans. Challenge vendor assumptions and ensure proposed solutions are technically sound, maintainable, secure, and aligned to client AI Platform Engineering standards. Define technical handover requirements, including source code, documentation, runbooks, test evidence, deployment instructions, and support procedures. Sign off vendor-delivered solutions from an engineering readiness perspective before production release. Ensure vendor-built components can be supported, reused, enhanced, and governed by client’s internal teams. WHATWE'RE LOOKING FOR Essential 7+ years of software or AI engineering experience, with at least 3 years in a senior technical lead or architect role Strong hands-on experience designing and delivering AI / ML / GenAI solutions in production environments. Deep understanding of AWS AI and cloud-native services, preferably including Bedrock, SageMaker, Textract, Comprehend, OpenSearch, Lambda, API Gateway, Step Functions, and S3. Strong practical knowledge of Generative AI engineering, including LLM integration, prompt engineering, RAG architecture, embeddings, vector stores, evaluation frameworks, and hallucination controls. Experience designing secure, scalable, observable, and maintainable production systems. Experience providing technical oversight of vendor-delivered AI solutions — reviewing designs, approving approaches, managing technical handover Production-grade Python — CI/CD, testing, observability, code review Line management or technical leadership experience — ability to lead, mentor and develop an engineering team Experience in a regulated industry — financial services or insurance strongly preferred Desirable AWS certifications: Solutions Architect Professional, Machine Learning Specialty Experience with Agentic AI, Bedrock Agents, Agent Core, or similar agent orchestration frameworks. Experience with LLMOps, MLOps, AI observability, model evaluation, model registry, and production model monitoring. Familiarity with AI risk management, model governance, responsible AI controls, and regulated data environments. Experience with vector databases or semantic search technologies such as OpenSearch, pgvector, Pinecone, Weaviate, or equivalent. Experience working across both classical ML product ionisation and modern GenAI solution delivery.